Method and device for video processing and medium
By introducing neural network architecture into video encoding and decoding technology, using parallel and serial processing methods, combining convolutional layer and activation layer, the problem of insufficient performance-complexity trade-off in the existing technology is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072687.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-13
- Publication Date
- 2025-05-23
AI Technical Summary
The existing video encoding and codec technology has shortcomings in performance-complexity trade-offs, making it difficult to further improve the encoding and codec efficiency.
A neural network architecture for video encoding and decoding is proposed. This architecture includes multiple branches for parallel processing and serial processing through multiple layers, combining convolutional layer and activation layer to improve the encoding and decoding performance.
Through this neural network architecture, performance-complexity trade-offs can be improved and codec performance can be improved.
Smart Images

Figure CN120035992A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to neural network architectures for video encoding. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] The embodiments of the present disclosure provide a solution for a neural network architecture for video encoding and decoding.
[0004] In a first aspect, a method for video processing is proposed. The method includes: obtaining a neural network (NN) model for processing a video, the NN model including at least one basic block, wherein the basic block includes: a plurality of branches for processing the input of the basic block in parallel, the branches including at least one convolutional layer and at least one activation layer, and a plurality of layers for serially processing the combination of the outputs of the plurality of branches, the plurality of layers including at least one convolutional layer and at least one activation layer; and performing conversion between a current video block of a video and a bit stream of the video according to the NN model. The method according to the first aspect of the present disclosure provides an efficient network architecture for video encoding and decoding, which can improve the performance-complexity trade-off. In this way, the encoding and decoding performance can be further improved.
[0005] In a second aspect, an apparatus for processing video data is provided, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to the first aspect.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, the instructions causing a processor to execute the method according to the first aspect.
[0007] In a fourth aspect, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream being generated by a method executed by a video processing device. The method includes: obtaining a neural network (NN) model for processing a video, the NN model including at least one basic block, wherein the basic block includes: a plurality of branches for processing the input of the basic block in parallel, the branches including at least one convolutional layer and at least one activation layer, and a plurality of layers for processing a combination of outputs of the plurality of branches in series, the plurality of layers including at least one convolutional layer and at least one activation layer; and generating a bitstream of the video according to the NN model.
[0008] In a fifth aspect, a method for storing a video bitstream is provided. The method includes: obtaining a neural network (NN) model for processing a video, the NN model including at least one basic block, wherein the basic block includes: multiple branches for processing the basic block's input in parallel, the branches including at least one convolutional layer and at least one activation layer, and multiple layers for serially processing a combination of the outputs of the multiple branches, the multiple layers including at least one convolutional layer and at least one activation layer; generating a video bitstream according to the NN model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 An example of raster scan striping of a picture is shown;
[0015] Figure 5 An example of rectangular strip segmentation of a picture is shown;
[0016] Figure 6 shows examples of pictures segmented into slices, tiles, and rectangular strips;
[0017] Figure 7A A schematic diagram showing a codec tree block (CTB) spanning a bottom picture boundary;
[0018] Figure 7B A schematic diagram of a CTB spanning the right picture boundary is shown;
[0019] Figure 7C A schematic diagram of a CTB spanning the bottom right picture boundary is shown;
[0020] Figure 8 An example of a VVC encoder block diagram is shown;
[0021] Figure 9 A schematic diagram showing picture samples on an 8×8 grid and horizontal and vertical block boundaries and non-overlapping blocks of 8×8 samples that can be deblocked in parallel;
[0022] Figure 10 Schematic diagram showing pixels involved in filter on / off decision and strong / weak filter selection;
[0023] Figure 11A An example of a 1-D directional pattern for EO point classification is shown, the 1-D directional pattern being a horizontal pattern with EO category = 0;
[0024] Figure 11B An example of a 1-D oriented pattern for EO point classification is shown, the 1-D oriented pattern being a vertical pattern with EO category = 1;
[0025] Figure 11C An example of a 1-D directional pattern for EO point classification is shown, which is a 135° diagonal pattern with EO category = 2;
[0026] Figure 11D An example of a 1-D directional pattern for EO point classification is shown, which is a 45° diagonal pattern with EO category = 3;
[0027] Figure 12A An example of a 5×5 diamond-shaped Geometric Transformation-Based Adaptive Loop Filter (GALF) filter shape is shown;
[0028] Figure 12B An example of a GALF filter shape of 7×7 diamond is shown;
[0029] Figure 12C An example of a 9×9 diamond GALF filter shape is shown;
[0030] Figure 13A An example of relative coordinates supported for a 5×5 diamond filter in the diagonal case is shown;
[0031] Figure 13B An example of relative coordinates supported for a 5×5 diamond filter in the case of vertical flipping is shown;
[0032] Figure 13C An example of relative coordinates supported for a 5×5 diamond filter in the case of rotation is shown;
[0033] Figure 14 An example of relative coordinates used for a 5×5 diamond filter support is shown;
[0034] Figure 15A A schematic diagram of the architecture of a commonly used convolutional neural network (CNN) is shown, where M represents the number of feature maps and N represents the number of samples in one dimension;
[0035] Figure 15B Shown Figure 15A An example of the construction of the residual block (ResBlock) in the CNN filter;
[0036] Figure 16A A schematic diagram illustrating the architecture of a first type (Type A) of basic blocks included in a NN model according to some embodiments of the present disclosure;
[0037] Figure 16B A schematic diagram illustrating the architecture of a second type (Type B) of basic blocks included in a NN model according to some embodiments of the present disclosure;
[0038] Figure 17 A schematic diagram showing the architecture of a NN model including three parts according to some embodiments of the present disclosure;
[0039] Figure 18 A schematic diagram showing a stack of basic blocks according to some embodiments of the present disclosure, wherein basic block type A is Figure 16A The block shown in ;
[0040] Figure 19 A schematic diagram showing a stack of basic blocks according to some further embodiments of the present disclosure, wherein the basic block type B is Figure 16B The block shown in ;
[0041] Figure 20 A schematic diagram showing a stack of basic blocks according to some further embodiments of the present disclosure, wherein basic block type A is Figure 16A The block shown in , and the basic block type B is Figure 16B The block shown in ;
[0042] Figure 21 A schematic diagram showing a stack of basic blocks according to some further embodiments of the present disclosure, wherein the basic block type B is Figure 16B The block shown in , and the basic block type A is Figure 16A The block shown in ;
[0043] Figure 22A A schematic diagram illustrating the architecture of a common residual block according to some embodiments of the present disclosure is shown;
[0044] Figure 22B A schematic diagram illustrating the architecture of a wide residual block according to some embodiments of the present disclosure, where M>K;
[0045] Figure 23 A schematic diagram illustrating the architecture of a proposed deep loop filter according to some embodiments of the present disclosure is shown;
[0046] Figure 24 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0047] Figure 25 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0048] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0049] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0050] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0051] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0052] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0053] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0054] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0055] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0056] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0057] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0058] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0059] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0060] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of FIG, video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0061] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0062] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0063] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0064] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0065] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0066] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0067] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0068] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0069] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0070] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0071] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0072] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0073] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0074] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0075] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0076] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0077] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0078] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0079] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0080] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0081] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0082] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0083] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0084] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0085] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0086] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0087] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0088] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0089] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0090] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0091] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Overview The present invention relates to video coding techniques. Specifically, the present invention relates to loop filters in image / video coding. The present invention can be applied to existing video coding standards, such as High Efficiency Video Codec (HEVC), Versatile Video Codec (VVC), or to-be-completed standards (e.g., AVS3). The present invention can also be applied to future video coding standards or video codecs, or used as a post-processing method outside the encoding / decoding process. 2. Background Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture in which temporal prediction and transform codecs are utilized. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard, with the goal of reducing bitrate by 50% compared to HEVC. VVC version 1 was completed in July 2020. 2.1. Color Space and Chroma Subsampling A color space (also called a color model (or color system)) is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually 3 or 4 values or color components (such as RGB). Basically, a color space is a refinement of coordinate systems and subspaces. For video compression, the most frequently used color spaces are YCbCr and RGB. YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, CB and CR are the blue difference and red difference chroma components. Y' (primed) is distinguished from Y, which is the luminance, meaning that light intensity is nonlinearly encoded based on the gamma-corrected RGB primaries. Chroma subsampling is the practice of encoding an image with less resolution for chroma information than for luma information, taking advantage of the human visual system's lower acuity for color differences than for luma. 2.1.1.4:4:4 Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2.4:2:2 The two chroma components are sampled at half the rate of luma: the horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by one-third with almost no visual difference. 2.1.3.4:2:0 In 4:2:0, horizontal sampling is doubled compared to 4:1:1, but because the Cb and Cr channels are sampled only on every alternate line in this scheme, the vertical resolution is halved. Therefore, the data rate is the same. Cb and Cr are each subsampled by a factor of 2 horizontally and vertically. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions. In MPEG-2, Cb and Cr are located at the same position horizontally, and Cb and Cr are located between pixels vertically (located in the gap). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in the gaps, midway between alternating luminance samples. In 4:2:0 DV, Cb and Cr are co-located horizontally and vertically on alternate lines. 2.2. Definition of Video Unit A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of a picture. A slice is divided into one or more bricks, each brick consisting of multiple CTU rows within the slice. A slice that is not split into multiple bricks is also called a brick. However, a brick that is a proper subset of a slice is not called a slice. A strip contains multiple slices of an image or multiple tiles of a slice. Two striping modes are supported: raster scan striping and rectangular striping. In raster scan striping, a strip contains a sequence of slices from a raster scan of the image. In rectangular striping, a strip contains multiple tiles that together form a rectangular region of the image. The tiles within a rectangular strip are in the order in which the tiles in the strip were raster scanned. Figure 4 An example of raster scan stripe partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan stripes. In this figure, a picture with 18×12 luma CTUs is partitioned into 12 slices and 3 raster scan stripes (informative). In the VVC specification Figure 5 An example of rectangular slice partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices. In this figure, a picture with 18×12 luma CTUs is partitioned into 24 slices and 9 rectangular slices (informative). In the VVC specification Figure 6 An example of a picture partitioned into slices, tiles, and strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular slices. In this figure, the picture is partitioned into 4 slices, 11 tiles, and 4 rectangular strips (informative). 2.2.1. CTU / CTB Size In VVC, the CTU size transmitted in the SPS by the syntax element log2_ctu_size_minus2 can be as small as 4×4. 7.3.2.3 Sequence Parameter Set RBSP Syntax log2_ctu_size_minus2 plus 2 specifies the luma codec tree block size for each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows: CtbLog2SizeY=log2_ctu_size_minus2+2 (7-9) CtbSizeY=1< <CtbLog2SizeY (7-10) MinCbLog2SizeY=log2_min_luma_coding_block_size_minus2+2(7-11) MinCbSizeY=1< <MinCbLog2SizeY (7-12) MinTbLog2SizeY = 2 (7 - 13) MaxTbLog2SizeY = 6 (7 - 14) MinTbSizeY = 1 << MinTbLog2SizeY (7 - 15) MaxTbSizeY = 1 << MaxTbLog2SizeY (7 - 16) PicWidthInCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY) (7 - 17) PicHeightInCtbsY = Ceil(pic_height_in_luma_samples ÷ CtbSizeY) (7 - 18) PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7 - 19) PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7 - 20) PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7 - 21) PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7 - 22) PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7 - 23) PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7 - 24) PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7 - 25) 2.2.2. CTUs in the Picture Assume a CTB / LCU size indicated by M×N (usually M equals N as defined in HEVC / VVC), and for a CTB located at the boundary of a picture (or slice or strip or other type, with the picture boundary taken as an example), K×L samples are within the picture boundary, where K < M or L < N. For CTBs such as Figure 7A , Figure 7B and Figure 7C depicted therein, the CTB size still equals M×N. However, the bottom boundary / right boundary of the CTB is outside the picture. Figure 7A , Figure 7B and Figure 7C show examples of CTBs spanning the picture boundary. Figure 7A shows a CTB spanning the bottom picture boundary, where K = M and L < N. Figure 7B shows a CTB spanning the right picture boundary, where K < M and L = N. Figure 7C shows a CTB spanning the bottom - right picture boundary, where K < M and L < N. 2.3. Encoding and Decoding Processes of Typical Video Codecs Figure 5 shows an example of an encoder block diagram of VVC, which includes three in - loop filter blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). Different from DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and by applying a Finite Impulse Response (FIR) filter respectively, where the encoded - decoded side information signals the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and fix artifacts created by previous stages. Figure 8 shows an example 800 of an encoder block diagram. 2.4. Deblocking Filter (DB) The input of DB is the reconstructed samples before the loop filter. The vertical edges in the picture are first filtered. Then, with the samples modified by the vertical - edge filtering process as the input, the horizontal edges in the picture are filtered. The vertical and horizontal edges in the CTB of each CTU are processed separately on a coding - decoding unit basis. The vertical edges in the coding - decoding blocks within the coding - decoding unit are filtered starting from the edge on the left - hand side of the coding - decoding block and advancing through the edges in the geometric order of the coding - decoding block towards the right - hand side. The horizontal edges in the coding - decoding blocks within the coding - decoding unit are filtered starting from the edge on the top of the coding - decoding block and advancing through the edges in the geometric order of the coding - decoding block towards the bottom. Figure 9An illustration of picture samples on an 8x8 grid along with horizontal and vertical block boundaries and non-overlapping blocks of 8x8 samples that can be deblocked in parallel is provided. 2.4.1. Boundary Decision The filter is applied at 8x8 block boundaries. Furthermore, the 8x8 block boundary must be a transform block boundary or a codec subblock boundary (e.g., due to the use of affine motion prediction (ATMVP)). For those boundaries that are not such, the filter is disabled. 2.4.2. Boundary strength calculation For transform block boundaries / codec sub-block boundaries, if the boundary is located in the 8×8 grid, the boundary can be filtered and the bS[xD i ][yD j ](where [xD i ][yD j ] represents coordinates) are defined in Table 1 and Table 2 respectively. Table 1. Boundary Strength (When SPS IBC is Disabled) Table 2. Boundary Strength (When SPS IBC is Enabled) 2.4.3. Deblocking Decision for Luma Component The deblocking decision is described in this subsection. Figure 10 The pixels involved in the filter on / off decision and strong / weak filter selection are shown. The wider, stronger luminance filter is a filter that is used only when conditions 1, 2, and 3 are all true. Condition 1 is the "large block condition". This condition detects whether the samples at the P side and Q side belong to a large block, which are represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows. bSidePisLargeBlk = (edge type is vertical and p0 belongs to a CU with width >= 32) || (edge type is horizontal and p0 belongs to a CU with height >= 32)? true:false bSideQisLargeBlk = ((edge type is vertical and q0 belongs to a CU with width >= 32) || (edge type is horizontal and q0 belongs to a CU with height >= 32))? True: False Based on bSidePisLargeBlk and bSideQisLargeBlk, Condition 1 is defined as follows. Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? True: False Next, if condition 1 is true, condition 2 will be checked further. First, the following variables are derived: Condition 2 = (d < β)? True: False Where d = dp0 + dq0 + dp3 + dq3. If conditions 1 and 2 are valid, then whether any block uses sub-blocks is further checked: Finally, if both conditions 1 and 2 are valid, the proposed deblocking method will check condition 3 (bulky strong filter condition), which is defined as follows. In the condition 3 strong filter condition, the following variables are derived: As in HEVC, strong filter conditions = (dpq less than (β>>2), sp3+sq3 less than (3*β>>5), and Abs(p0-q0) less than (5*t C +1)>>1)? True:False. 2.4.4. Strong deblocking filter for luminance (designed for larger blocks) When samples on either side of the boundary belong to a large block, a bilinear filter is used. Samples belonging to a large block are defined as when width >= 32 for vertical edges and when height >= 32 for horizontal edges. Bilinear filters are listed below. For block boundary samples p from i=0 to Sp-1 i and block boundary samples q for j = 0 to Sq-1 i (pi and qi are the i-th sample points in the row for filtering vertical edges, or the i-th sample points in the column for filtering horizontal edges) are then replaced by linear interpolation as follows: —p i ′=(f i *Middle s,t +(64-f i )*P s +32)>>6), is clipped to p i ±tcPD i —q j ′=(gi *Middle s,t +(64-g j )*Q s +32)>>6), is limited to q j ±tcPD j Among them, tcPD i and tcPD j The term is the position-dependent clipping described in Section 2.4.7, and g j ,f i ,Middle s,t ,P s and Q s is given below. 2.4.5. Deblocking Control for Chroma A chroma strong filter is applied on both sides of the block boundary. Here, the chroma filter is selected when the distance between the two sides of the chroma edge is greater than or equal to 8 (chroma position), and the following decisions are met with three conditions: The first condition is used for boundary strength and large block decisions. The proposed filter can be applied when the block width or block height orthogonally crossing the block edge is equal to or greater than 8 in the chroma sample domain. The second and third conditions are essentially the same for the HEVC luma deblocking decision, which is an on / off decision and a strong filter decision, respectively. In the first decision, the boundary strength (bS) is modified for chroma filtering and the condition is then checked. If the condition is met, the remaining conditions with lower priority are skipped. When bS is equal to 2, chroma deblocking is performed, or bS is equal to 1 when a large block boundary is detected. The second and third conditions are essentially the same as the HEVC luma strong filter decision as follows. In the second condition: d is then derived as in HEVC luma deblocking. When d is less than β, the second condition will be true. In the third condition, the strong filter condition is derived as follows: dpq is derived as in HEVC. sp3 = Abs(p3 - p0), as derived in HEVC. sq3=Abs(q0-q3), as derived in HEVC. As in HEVC design, strong filter conditions = (dpq less than (β>>2), sp3+sq3 less than (β>>3), and Abs(p0-q0) less than (5*t C +1)>>1). 2.4.6. Strong Deblocking Filter for Chroma The following strong deblocking filter for chroma is defined: p2′=(3*p3+2*p2+p1+p0+q0+4)>>3 p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3 p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3 The proposed chroma filter performs deblocking on a 4×4 grid of chroma samples. 2.4.7. Position-dependent clipping Position-dependent clipping tcPD is applied to the output samples of the luma filtering process, which involves modifying strong and long filters of 7, 5, and 3 samples at the boundaries. Assuming the quantization error distribution, it is recommended to increase the clipping value for samples that are expected to have higher quantization noise, and therefore higher deviations of the reconstructed sample values from the true sample values. For each P-edge or Q-edge filtered with the asymmetric filter, a position-dependent threshold table is selected from the two tables (i.e., Tc7 and Tc3 listed below) provided to the decoder as side information, depending on the result of the decision-making process in Section 2.4.2: Tc7={6,5,4,3,2,1,1}; Tc3={6,4,2}; tcPD=(Sp==3)? Tc3:Tc7; tcQD=(Sq==3)? Tc3:Tc7; For P-edges or Q-edges filtered with a short symmetric filter, a lower magnitude position-dependent threshold is applied: Tc3={3,2,1}; After defining the threshold, the filtered p' i and q' i Sample values are clipped: p” i =Clip3(p' i +tcP i ,p' i –tcP i ,p' i ); q” j =Clip3(q' j +tcQ j ,q' j –tcQ j ,q' j ); where p' i and q' i is the filtered sample value, p” i and q” j is the output sample value after limiting, and tcP i tcP i is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. Function Clip3 is the clipping function as specified in VVC. 2.4.8. Sub-block Deblocking Adjustment To enable parallel-friendly deblocking using long filters and subblock deblocking, the long filter is restricted to modifying at most 5 samples on a side using subblock deblocking (affine or ATMVP or DMVR) as shown in the luma control for the long filter. In addition, subblock deblocking is adjusted so that subblock boundaries on the 8×8 grid close to CU or implicit TU boundaries are restricted to modifying at most two samples on each side. The following applies to sub-block boundaries that are not aligned with CU boundaries. Where edge equals 0 corresponds to the CU boundary, edge equals 2 or equal to orthogonalLength-2 corresponds to the sub-block boundary 8 samples from the CU boundary, etc. If implicit partitioning of TU is used, implicitTU is true. 2.5.SAO The input of SAO is the reconstructed samples after DB. The concept of SAO is to reduce the average sample distortion of the region by first classifying the region samples into multiple categories with a selected classifier, obtaining an offset for each category, and then adding the offset to each sample of the category, where the classifier index and the offset of the region are encoded in the bitstream. In HEVC and VVC, the region (the unit for SAO parameter signaling) is defined as CTU. Two SAO types that can meet the low complexity requirements are adopted in HEVC. These two types are edge offset (EO) and band offset (BO), which will be discussed in further detail below. The index of the SAO type is encoded (it is in the range of [0, 2]). For EO, the sample classification is based on the comparison between the current sample and the neighboring samples according to the one-dimensional strip direction mode (horizontal, vertical, 135° diagonal and 45° diagonal). 11A to 11D Four 1-D band orientation patterns are shown for EO point classification: horizontal (EO category = 0), vertical (EO category = 1), 135° diagonal (EO category = 2), and 45° diagonal (EO category = 3). For a given EO category, each sample point within the CTB is classified into one of five categories. The current sample point value, which is labeled "c", is compared with its two nearest neighbors along the selected 1-D pattern. The classification rules for each sample point are summarized in Table 1. Category 1 and Category 4 are associated with local valleys and local peaks, respectively, along the selected 1-D pattern. Category 2 and Category 3 are associated with concave corners and convex corners, respectively, along the selected 1-D pattern. If the current sample point does not belong to EO category 1 to 4, the current sample point is category 0 and SAO is not applied. Table 3: Sample classification rules for edge compensation type condition 1 c < a and c < b 2 (c<a&&c==b)||(c==a&&c<b) 3 (c>a&&c==b)||(c==a&&c>b) 4 c>a&&c>b 5 None of the above 2.6. Geometric Transformation-Based Adaptive Loop Filter in JEM The input of DB is the reconstructed samples after DB and SAO. The sample classification and filtering process is based on the reconstructed samples after DB and SAO. In JEM, a geometric transformation-based adaptive loop filter (GALF) with block-based filter adaptation [3] is applied. For the luminance component, one of 25 filters is selected for each 2×2 block based on the direction and activity of the local gradient. Filter shape In JEM, up to three diamond filter shapes (such as Figures 12A-12C (as shown) can be selected for the luma component. An index is signaled at the picture level to indicate the filter shape to be used for the luma component. Each square represents a sample, and Ci (i is 0 to 6 (left), 0 to 12 (center), 0 to 20 (right)) represents the coefficient to be applied to the sample. For the chroma components in a picture, a 5×5 diamond shape is always used. Figures 12A-12C GALF filter shapes are shown (left: 5x5 diamond, middle: 7x7 diamond, right: 9x9 diamond). 2.6.1.1. Block Classification Each 2×2 block is classified into one of 25 categories. The classification index C is based on the directionality D and activity of the 2×2 block. The quantized value of is derived as follows: To calculate D and The horizontal, vertical and two diagonal gradients are first calculated using the 1-D Laplacian operator: The indices i and j refer to the coordinates of the top left sample in the 2x2 block and R(i, j) indicates the reconstructed sample at coordinates (i, j). Then the maximum and minimum values of D for the horizontal and vertical gradients are set as: And the maximum and minimum values of the gradients in the two diagonal directions are set to: To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2: Step 1. If and is true, D is set to 0. Step 2. If Then continue with step 3; otherwise, continue with step 4. Step 3. If Then D is set to 2; otherwise D is set to 1. Step 4. If Then D is set to 4; otherwise D is set to 3. The activity value A is calculated as: A is further quantized into a range of 0 to 4 (including the boundary values), and the quantized value is represented as For the two chroma components in a picture, no classification method is applied, ie, a single set of ALF coefficients is applied for each chroma component. 2.6.1.2. Geometric transformation of filter coefficients Figure 13A The relative coordinates for a 5x5 diamond filter support in the diagonal case are shown. Figure 13B The relative coordinates supported for a 5x5 diamond filter in the case of vertical flipping are shown. Figure 13C The relative coordinates for a 5x5 diamond filter support under rotation are shown. Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients associated with the coordinates (k, l), depending on the gradient values calculated for the block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to create different blocks by aligning the directionality of the ALF, in which the ALF is applied. Three geometric transformations, including diagonal, vertical flip and rotation, are introduced: Diagonal:f D (k,l)=f(l,k), Flip vertically:f V (k,l)=f(k,Kl-1), (9) Rotation: f R (k,l)=f(Kl-1,k). Where K is the size of the filter and 0≤k,l≤K-1 are the coefficient coordinates, such that position (0,0) is in the upper left corner and position (K-1,K-1) is in the lower right corner. Depending on the gradient value calculated for the block, a transform is applied to the filter coefficients f(k,l). The relationship between the transform and the four gradients in the four directions is summarized in Table 4. Figures 13A-13C The transform coefficients for each position are shown based on a 5x5 diamond. Table 4: Mapping and transformation of gradients computed for a block Gradient value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> diagonal <![CDATA[g d1 <g d2 And g h <g v ]]> Flip vertically <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation 2.6.1.3. Filter parameter signaling In JEM, GALF filter parameters are signaled for the first CTU (i.e., after the slice header and before the SAO parameters of the first CTU). A set of up to 25 luma filter coefficients can be signaled. To reduce bit overhead, filter coefficients of different categories can be merged. In addition, the GALF coefficients of the reference picture are stored and allowed to be used as the GALF coefficients of the current picture. The current picture can choose to use the GALF coefficients stored for the reference picture and bypass GALF coefficient signaling. In this case, only the index of one of the reference pictures is signaled, and the stored GALF coefficients of the indicated reference picture are inherited for the current picture. To support GALF temporal prediction, a candidate list of GALF filter sets is maintained. At the start of decoding a new sequence, the candidate list is empty. After decoding a picture, the corresponding set of filters can be added to the candidate list. Once the candidate list size reaches the maximum allowed value (i.e., 6 in the current JEM), the new set of filters overwrites the oldest set in decoding order, and in other words, the first-in-first-out (FIFO) rule is applied to update the candidate list. To avoid duplication, a set is only added to the list if the corresponding picture does not use GALF temporal prediction. To support temporal scalability, multiple candidate lists of filter sets exist, and each candidate list is associated with a temporal layer. More specifically, each array assigned by a temporal layer index (TempIdx) can be composed of filter sets from the previously decoded picture with the same lower TempIdx. For example, the kth array is assigned to be associated with a TempIdx equal to k, and this array only contains filter sets from pictures with a TempIdx less than or equal to k. After encoding and decoding a picture, the filter set associated with the picture will be used to update those arrays associated with equal or higher TempIdx. Temporal prediction of the GALF coefficients is used for inter-frame coding to minimize signaling overhead. For intra frames, temporal prediction is not available, and a set of 16 fixed filters is assigned to each class. To indicate the use of fixed filters, a flag for each class is signaled, and if necessary, the index of the selected fixed filter is signaled. Even when a fixed filter is selected for a given class, the coefficients f(k, l) of the adaptive filter can still be sent for that class, in which case the coefficients of the filter to be applied in the reconstructed image are the sum of the two sets of coefficients. The filtering process of the luma component can be controlled at the CU level. A flag is signaled to indicate whether GALF will be applied to the luma component of a CU. For chroma components, whether GALF will be applied is indicated only at the picture level. 2.6.1.4. Filtering process At the decoder side, when GALF is enabled for a block, each sample R(i,j) within the block is filtered, resulting in a sample value R′(i,j) as shown below, where L represents the filter length and f m,n denotes the filter coefficients, and f(k, l) denotes the decoding filter coefficients. Figure 14 An example of relative coordinates used for a 5×5 diamond filter support is shown, which assumes that the coordinates (i, j) of the current sample point are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficients. Figure 14 In , examples of relative coordinates supported for a 5×5 diamond filter are provided. 2.7. Geometric Adaptive Loop Filter (GALF) in VVC 2.7.1.GALF in VTM-4 In VTM4.0, the filtering process of the adaptive loop filter is performed as follows: O(x,y)=∑ (i,j) w(i,j).I(x+i,y+j), (11) Where the sample point I(x+i,y+j) is the input sample point, O(x,y) is the filtered output sample point (i.e., the filter result), and w(i,j) represents the filter coefficient. In practice, in VTM4.0, integer arithmetic operations are used to implement fixed-point precision calculations: where L represents the filter length, and where w(i,j) are the filter coefficients in fixed-point precision. Compared with the GALF in JEM, the current design of the GALF in VVC has the following major changes: 1) Adaptive filter shapes are removed. Only 7x7 filter shapes are allowed for luma components and 5x5 filter shapes are allowed for chroma components. 2) Signaling of ALF parameters from slice / picture level to CTU level is removed. 3) Class index calculation is performed at the 4×4 level instead of the 2×2 level. Furthermore, a subsampled Laplacian calculation method for ALF classification is utilized, as proposed in JVET-L0147. More specifically, the horizontal / vertical / 45° / 135° gradient for each sample point within a block does not need to be calculated. Instead, 1:2 subsampling is utilized. 2.8. Nonlinear ALF in Current VVC 2.8.1. Filtered Representation Without the influence of codec efficiency, equation (11) can be re-expressed in the following expression: O(x,y)=I(x,y)+∑ (i,j)≠(0,0) w(i,j).(I(x+i,y+j)-I(x,y)), (13) where w(i,j) are the same filter coefficients as in (11) [except for w(0,0), which is equal to 1 in (13) and 1-∑ (i,j)≠(0,0) w(i,j)]. Using the above filter formula of (13), VVC introduces nonlinearity to make the ALF more effective by using a simple clipping function to reduce the influence of neighboring sample values (when (I(x+i,y+j)) is too different from the current sample value I(x,y)), I(x+i,y+j)) is filtered). More specifically, the ALF filter is modified as follows: O′(x,y)=I(x,y)+∑ (i,j)≠(0,0) w(i,j).k(I(x+i,y+j)-I(x,y),k(i,j)), (14) where K(d,b)=min(b,max(-b,d)) is the clipping function and k(i,j) is the clipping parameter that depends on the (i,j) filter coefficients. The codec performs an optimization to find the best k(i,j). In the JVET-N0242 implementation, clipping parameters are specified for each ALF filter, and for each filter coefficient, one clipping value is signaled. This means that for each luma filter, up to 12 clipping values can be signaled in the bitstream, and for each chroma filter, up to 6 clipping values can be signaled in the bitstream. In order to limit the signaling cost and codec complexity, only 4 fixed values are used, which are the same for inter and intra slices. Because the variance of local differences for luma is typically higher than that for chroma, two different sets of luma filters and chroma filters are applied. A maximum sample value in each set (here 1024 for 10-bit bit depth) is also introduced so that clipping can be disabled if it is not necessary. The set of clipping values used in the JVET-N0242 test is provided in Table 5. The 4 values have been chosen by roughly equally dividing the full range of sample values for luma (encoded on 10 bits) and the range of 4 to 1024 for chroma in the logarithmic domain. More precisely, the brightness table of the clipping values has been obtained by the following formula: Where M = 2 10 And N=4. (15) Similarly, the chromaticity table of the clipped value is obtained according to the following formula: Where M = 2 10 , N = 4 and A = 4. (16) Table 5: Authorized limit values The selected clipping value is encoded in the "alf_data" syntax element using the Golomb coding scheme corresponding to the index of the clipping value in the above Table 5. This coding scheme is the same as that for the filter index. 2.9. Convolutional Neural Network-Based Loop Filters for Video Codecs Convolutional Neural Networks In deep learning, convolutional neural networks (CNNs or ConvNets) are a type of deep neural network that is most commonly used to analyze visual images. Convolutional neural networks have been very successful in image and video recognition / processing, recommendation systems, image classification, medical image classification, and natural language processing. CNNs are regularized versions of multilayer perceptrons. Multilayer perceptrons are usually referred to as fully connected networks, meaning that every neuron in one layer is connected to all neurons in the next layer. The "full connectivity" of these networks makes them prone to overfitting the data. Typical approaches to regularization involve adding some form of magnitude measure of the weights to the loss function. CNNs take a different approach to regularization: they exploit hierarchical patterns in the data and use smaller and simpler patterns to assemble more complex patterns. Therefore, on the scale of connectivity and complexity, CNNs are on the lower end. Compared to other image classification / processing algorithms, CNNs use relatively little preprocessing. This means the network learns filters that are manually engineered in traditional algorithms. This independence from prior knowledge and human effort in feature design is a major advantage. 2.9.2. Deep Learning for Image / Video Encoding and Decoding Image / video compression based on deep learning generally has two meanings: pure neural network-based end-to-end compression and traditional frameworks enhanced by neural networks. The first type generally adopts an autoencoder-like structure implemented by a convolutional neural network or a recurrent neural network. Although relying solely on neural networks for image / video compression can avoid any manual optimization or hand-design, the compression efficiency may be unsatisfactory. Therefore, work with the second type of distribution uses neural networks as an auxiliary and enhances the traditional compression framework by replacing or enhancing some modules. In this way, work with the second type of distribution can inherit the advantages of highly optimized traditional frameworks. For example, Li et al. proposed a fully connected network for intra-frame prediction in HEVC. In addition to intra-frame prediction, deep learning has also been developed to enhance other modules. For example, Dai et al. used a convolutional neural network to replace the in-loop filter of HEVC and achieved promising results. This work applied neural networks to improve the arithmetic codec engine. 2.9.3. In-loop filtering based on convolutional neural networks In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, resulting in distortion of the reconstructed frame. To mitigate this distortion, a convolutional neural network can be trained to learn the mapping from the distorted frame to the original frame. In practice, training must be performed before deploying the CNN-based loop filter. 2.9.3.1. Training The goal of the training process is to find optimal values for the parameters including weights and biases. First, a codec (e.g., HM, JEM, VTM, etc.) is used to compress the training dataset to generate distorted reconstructed frames. The reconstructed frame is then fed into the CNN, and the cost is calculated using the output of the CNN and the real frame (original frame). Common cost functions include SAD (sum of absolute differences) and MSE (mean square error). Next, the gradient of the cost for each parameter is derived through the backpropagation algorithm. Using the gradient, the value of the parameter can be updated. The above process is repeated until the convergence criterion is met. After training is completed, the derived optimal parameters are saved for use in the inference phase. 2.9.3.2 Convolution Processing During convolution, the filter moves across the image from left to right and top to bottom, with a 1-pixel column change for horizontal movement and a 1-pixel row change for vertical movement. The amount of movement between applications of the filter to the input image is called the stride, and it is almost always symmetric across the height and width dimensions. For both height and width movement, the default stride or stride is (1,1) in both dimensions. In most deep convolutional neural networks, residual blocks are used as basic modules and are stacked several times to build the final network, where in one example, the residual blocks are obtained by combining convolutional layers, ReLU / PReLU activation functions and convolutional layers, as shown in Figure 15A and Figure 15B shown. Figure 15A The following figure shows the architecture of a commonly used CNN. M represents the number of feature maps. N represents the number of samples in one dimension. Figure 15B Shown Figure 15A The construction of ResBlock (residual block) in . 2.9.3.3. Reasoning During the inference phase, the distorted reconstructed frame is fed into the CNN and processed by the CNN model, whose parameters have been determined in the training phase. The input samples of the CNN can be reconstructed before or after DB, or before or after SAO, or before or after ALF. 3. Question Current neural network-based encoding and decoding tools have the following problems: 1. The performance-complexity trade-off needs to be further improved. For example, higher codec gains are achieved at the cost of a neural network with higher computational complexity. To achieve a better trade-off, more efficient network architectures should be investigated. 2. Most video codecs are based on network architectures inherited from computer vision tasks such as image classification, object detection, or image restoration tasks such as image super-resolution, image denoising, etc. 4. Examples The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. One or more neural network (NN) models are trained as codec tools to improve the efficiency of video codecs. These NN-based codec tools can be used to replace or enhance modules involved in video codecs. For example, the NN model can be used as an additional intra-frame prediction mode, inter-frame prediction mode, transform kernel, or loop filter. These embodiments detail how to design the NN model by using external information such as prediction, partitioning, QP, etc. It should be noted that the NN model can be used as any codec tool, such as NN-based intra / inter prediction, NN-based super-resolution, NN-based motion compensation, NN-based reference generation, NN-based fractional pixel interpolation, NN-based loop / post-processing filtering, etc. In the present disclosure, the NN model can be any kind of NN architecture, such as a convolutional neural network or a fully connected neural network, or a combination of a convolutional neural network and a fully connected neural network. In the following discussion, a video unit can be a sequence, a picture, a slice, a tile, a sub-picture, a CTU / CTB, a CTU row / CTB row, one or more CUs / CBs, one or more CTUs / CTBs, one or more VPDUs (virtual pipeline data units), or a sub-region within a picture / slice / slice / tile. A parent video unit represents a unit larger than a video unit. Typically, a parent unit will contain several video units. For example, when the video unit is a CTU, the parent unit can be a slice, a CTU row, multiple CTUs, and so on. About a New Basic Block 1. Figure 16A and / or Figure 16B The block shown in and / or any variation of this block can be used as a basic block for building a NN model. 16A to 16B The basic blocks included in the NN model are shown. a. In one example, 16A to 16B In the figure, the regular rectangles represent convolutional layers (Layer1, Layer2, Layer5, Layer7), while the rounded rectangles represent activation layers (Layer3, Layer4, Layer6, Layer8). The arrows show the data flow. According to the direction of the arrow, the output of the previous layer is used as the input of the next layer. out ,C in ,K hor ,K ver,S refers to the number of output channels, the number of input channels, the kernel size in the horizontal direction, the kernel size in the vertical direction, and the stride of the convolutional layer, respectively. Compared with (a), (b) includes a skip connection at the end, that is, the input of Layer1 (which is also the input of Layer2) is added to the output of Layer8. b. Where the output of a preceding layer is fed into multiple next layers, the input of each next layer is the same as the output. c. In the case where the outputs of multiple previous layers are fed into the next layer, the input to the next layer is the concatenation of these outputs along the channel dimension. 2. 16A to 16B The activation layers shown in can be configured in any flexible way. a. In one example, Figure 16A and / or Figure 16B At least one of the activation layers in is a nonlinear function. b. In one example, Figure 16A and / or Figure 16B At least one of the activation layers in is a linear function. c. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is PReLU (parameterized rectified linear unit) d. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e., f1 and / or f2) in is LReLU (leaky rectified linear unit). e. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is ReLU (rectified linear unit). f. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is the identity mapping function (the output and input of this layer are exactly the same). g. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e. f3 and / or f4) in is PReLU (parameterized rectified linear unit). h. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is LReLU (leaky rectified linear unit). i. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e. f3 and / or f4) in is ReLU (rectified linear unit). j. In one example, Figure 16A and / or Figure 16B All activation layers in
[15] are PReLU (parameterized rectified linear units). k. In one example, Figure 16A and / or Figure 16B All activation layers in
[15] are LReLU (leaky rectified linear units). 1. In one example, Figure 16A and / or Figure 16B All activation layers in are ReLU (Rectified Linear Units). m. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is a nonlinear function, and Figure 16A and / or Figure 16B Layer6 and / or Layer8 (ie, f3 and / or f4) in are linear functions. n.In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is a nonlinear function, and Figure 16A and / or Figure 16B Layer6 and / or Layer8 (ie, f3 and / or f4) in are identity mapping functions. i. In one example, Figure 16A and / or Figure 16B Lyaer3 and Lyaer4 (i.e. f1 and f2) in are PReLU (parameterized rectified linear units), and Figure 16A and Figure 16B Layer6 and Layer8 (i.e. f3 and f4) in are identity mapping functions. o. In one example, which configuration to use can be determined by decoding information. p. In one example, which configuration to use may be determined by at least one syntax element signaled from the encoder to the decoder. 3. 16A to 16B The convolutional layers shown in can be configured in any flexible way. a. In one example, the kernel size in all layers is the same, e.g., 1×1, 3×3, 5×5, etc. b. In one example, the kernel sizes in all layers are different. c. In one example, the kernel sizes in Layer 1 and Layer 2 are different, while the kernel sizes in Layer 5 and Layer 7 are the same. i. In one example, the kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3×3. ii. In one example, the kernel sizes in Layer 1 and Layer 2 are 1×1 and 5×5 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3×3. iii. In one example, the kernel sizes in Layer 1 and Layer 2 are 3×3 and 1×1 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3×3. iv. In one example, the kernel sizes in Layer 1 and Layer 2 are 5×5 and 1×1 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3×3. d. In one example, the kernel sizes in Layer 1 and Layer 2 are different, and the kernel sizes in Layer 5 and Layer 7 are also different. i. In one example, the kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3 respectively, while the kernel sizes in Layer 5 and Layer 7 are 1×1 and 3×3 respectively. ii. In one example, the kernel sizes in Lyaer1 and Lyaer2 are 3×3 and 1×1 respectively, while the kernel sizes in Layer5 and Layer7 are 1×1 and 3×3 respectively. e. In one example, the number of output channels (also called output feature maps) in all layers is the same, for example, 32, 64, 96, 128, etc. f. In one example, the number of output channels in all layers is different. g. In one example, the number of output channels in Layer 1 and Layer 2 is different, while the number of output channels in Layer 5 and Layer 7 is the same. i. In one example, the number of output channels in Layer 1 is greater than the number of output channels in Layer 2. ii. In one example, the number of output channels in Layer 1 is smaller than the number of output channels in Layer 2. h. In one example, the number of output channels in Lyaer1 and Lyaer2 is different, and the number of output channels in Lyaer5 and Lyaer7 is also different. i. In one example, the number of output channels in Lyaer1 is greater than the number of output channels in Lyaer2, and the number of output channels in Lyaer5 is greater than the number of output channels in Lyaer7. ii. In one example, the number of output channels in Layer 1 is smaller than the number of output channels in Layer 2, and the number of output channels in Layer 5 is larger than the number of output channels in Layer 7. i. In one example, the kernel sizes in Layer 1 and Layer 2 are different, the kernel sizes in Layer 5 and Layer 7 are different, the number of output channels in Layer 1 and Layer 2 are different, and the number of output channels in Layer 5 and Layer 7 are the same. i. In one example, let the number of input channels in Layer 1 (which is also the input channel of Layer 2) be N, and the number of output channels in Layer 1 and Layer 2 be set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 (e.g., 2.5) and C2 is less than 1.0 (e.g., 0.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively. ii. In one example, let the number of input channels of Layer 1 (which is also the input channel of Layer 2) be N, and the number of output channels in Layer 1 and Layer 2 be set to N×C1 and N×C2, respectively, where C1 is less than 1.0 (e.g., 0.5) and C2 is greater than 1.0 (e.g., 2.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively. iii. In one example, given an integer N (e.g., 32, 64, 96, 128, etc.), the number of output channels in Layer 1 and Layer 2 is set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 (e.g., 2.5) and C2 is less than 1.0 (e.g., 0.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively. iv. In one example, given an integer N (e.g., 32, 64, 96, 128, etc.), the number of output channels in Layer 1 and Layer 2 is set to N×C1 and N×C2, respectively, where C1 is less than 1.0 (e.g., 0.5) and C2 is greater than 1.0 (e.g., 2.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively. j. In one example, which configuration to use can be determined by decoding the information. k. In one example, which configuration to use may be determined by at least one syntax element signaled from the encoder to the decoder. 4. The configurations in the above two items (item 2 and item 3) can be combined. Specifically, the activation layer can be configured using any sub-item (2.a, 2.b, ..., 2.p) from item 2, and the convolution layer can be configured using any sub-item (3.a, 3.b, ..., 3.k) from item 3. 5. 16A to 16B The activation layers and convolutional layers shown can be jointly configured in certain ways to achieve better performance. a. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is a nonlinear function, and Figure 16A and / or Figure 16B Layer 6 and / or Layer 8 (i.e., f3 and / or f4) in
[15] are identity mapping functions. The kernel sizes in Layer 1 and Layer 2 are different, the kernel sizes in Layer 5 and Layer 7 are different, the number of output channels in Layer 1 and Layer 2 is different, and the number of output channels in Layer 5 and Layer 7 is the same. i. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e. f1 and / or f2) in is PReLU (parameterized rectified linear unit) ii. In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e., f1 and / or f2) in is LReLU (leaky rectified linear unit). iii. In one example, Figure 16A and / or Figure 16BLyaer3 and / or Lyaer4 (i.e. f1 and / or f2) in is ReLU (rectified linear unit). iv. In one example, the kernel size in Lyaer1 is smaller than that in Lyaer2, the kernel size in Lyaer5 is smaller than that in Layer7, and the number of output channels in Layer1 is greater than that in Layer2. v. In one example, the kernel size in Layer 1 is larger than the kernel size in Layer 2, the kernel size in Layer 5 is smaller than the kernel size in Layer 7, and the number of output channels in Layer 1 is smaller than the number of output channels in Layer 2. About NN Model 6.NN models can include 16A to 16B One or more basic blocks shown in . a. In one example, the NN model includes Figure 16A At least one basic block is shown. b. In one example, the NN model includes Figure 16B At least one basic block is shown. c. In one example, the NN model includes Figure 16A At least one basic block shown and Figure 16B At least one basic block is shown. 7.NN models can include Figure 17 The three parts shown in the figure are as follows: the head part is responsible for extracting features from the input of the NN model, which are then fed into the backbone part for further feature mapping. The tail part transforms the output features of the backbone part into the final output. a. In one example, the header section begins with Figure 18 The way shown is designed, Figure 18 shows a stack of basic blocks, where basic block type A is Figure 16A The blocks shown, and Ma is the number of stacked blocks. b. In one example, the header section begins with Figure 19 The way shown is designed, Figure 19 shows a stack of basic blocks, where basic block type B is Figure 16B The blocks shown, and Mb is the number of stacked blocks. c. In one example, the header section begins with Figure 20 The way shown is designed, Figure 20 shows a stack of basic blocks, where basic block type A is Figure 16A The block shown, and the basic block type B is Figure 16B The block shown, M a and M bis the number of blocks stacked. d. In one example, the header begins with Figure 21 The way it is designed, Figure 21 shows a stack of basic blocks, where basic block type B is Figure 16B The blocks shown, and the basic block type A is Figure 16A The block shown in M a and N b is the number of blocks stacked. e. In one example, the backbone portion is Figure 18 The manner shown is designed. f. In one example, the backbone portion is Figure 19 The manner shown is designed. g. In one example, the backbone portion is Figure 20 The manner shown is designed. h. In one example, the backbone portion is Figure 21 The manner shown is designed. i. In one example, the tail portion begins with Figure 18 The manner shown is designed. j. In one example, the tail portion begins with Figure 19 The manner shown is designed. k. In one example, the tail portion begins with Figure 20 The manner shown is designed. l. In one example, the tail portion begins with Figure 21 The manner shown is designed. m. In one example, the header begins with Figure 18 The backbone is designed in the manner shown. Figure 21 The tail part is a conventional convolutional layer. 8. In one example, only integer operations can be applied in the proposed architecture. a. Floating-point operations may not be involved. b. Division operations may not be involved. c. Integer operations may include "addition", "multiplication", "shift", "rounding", "limiting", etc. 5. Example Implementation 5.1. Summary This paper proposes a deep loop filter based on a basic residual block with wide activation and large receptive field. The proposed filter is implemented on top of the NNVC general software NCS-1.0. The BD rate changes of {Y, Cb, Cr} on NCS-1.0 and NNVC-2.0 are summarized as follows: Based on NCS-1.0: Conventional: RA: {%,%,%,}, LB: {%,%,%,}, AI: {-1.55%,-1.94%,-2.12%} Compact: RA:{%,%,%,%}, LB:{%,%,%,%}, AI:{%,%,%,%} Based on: NNVC-2.0: Conventional: RA: {%,%,%,}, LB: {%,%,%,}, AI: {-8.68%,-21.49%,-22.09%} Compact: RA:{%,%,%,%}, LB:{%,%,%,%}, AI:{%,%,%,%} 5.2. Introduction The NNVC general software includes two sets of deep in-loop filters, where filter set 1 is based on Figure 22A The residual block shown in Example 2200A of FIG. 2 includes two 3×3 convolutional layers, which is a normal residual block. Figure 22B The benefit of building a network with a residual block with wide activation as shown in Example 2200B of
[15] is that the residual block is a wide residual block with M>K. However, compared with ordinary residual blocks, the receptive field of wide residual blocks is limited. This manuscript proposes a new type of residual block with wide activation and large receptive field, such as Figure 23 The architecture is shown in 2300. In addition, the proposed residual block allows multi-scale feature extraction. 5.3. Proposed method Sections 2.1 to 2.4 present the luma CNN architecture, chroma CNN architecture, inference, and training procedures, respectively. Note that other designs of the proposed method (such as parameter selection, residual scaling, combination with deblocking, etc.) remain the same as those of NN-based filter set 1 in NCS-1.0. 5.3.1. Brightness CNN Structure Figure 23 The architecture of the proposed CNN filter for deep loop filtering is given, which consists of three types of basic blocks called head blocks, backbone blocks and tail blocks. The design of these blocks follows the principles of wide activation, large receptive field and multi-scale feature extraction. The head block is responsible for extracting features from the input, C inrepresents the number of input channels and is equal to 5 for the intra model (reconstruction, prediction, partitioning, block mode signal, quantization parameter) and equal to 3 for the inter model (reconstruction, prediction, quantization parameter). C represents the basic number of feature maps and is set to 64. {C1, C2} represents the number of output channels in the large activation branch and the large receptive field branch and is set to {160, 32}. C refers to the stride of the convolution and is set to 2 to achieve feature downsampling. The backbone of the proposed network, which contains a series of backbone blocks, implements feature embedding. The number of backbone blocks N is set to 22 and 19 for regular models and compact models. At the end, there is a tail block that maps the embedded features from the backbone to the final output. 5.3.2. Chroma CNN Architecture The CNN filter for chroma components has a similar architecture to that for luma, but includes fewer backbone blocks. Specifically, N is set to 10. 5.3.3. Reasoning SADL is used to perform inference of the proposed CNN filters in VTM. Both weights and feature maps are represented in int16 precision using static quantization. As suggested, the network information in the inference phase is provided in Table 6. Table 6. Network information for NN-based video codec tool testing during the inference phase 5.3.4. Training PyTorch was used as the training platform. The DIV2K and BVI-DVC datasets were used to train the CNN filters for I-slice and B-slice, respectively. As suggested, the network information during the training phase is provided in Table 7. Table 7. Network information used for NN-based video codec tool testing during the training phase 5.4 Experimental results The proposed CNN-based loop filtering method was integrated into NCS-1.0 and tested according to common test conditions. After the proposed CNN-based filtering, SAO was disabled and ALF (and CCALF) was placed. In order to better evaluate the proposed method, the proposed method was compared with the NN-based filter set 1 of NCS-1.0 and NNVC-2.0. The comparison results are shown in Tables 8 to 13. Conventional model: The proposed model with a regular size includes 22 backbone blocks. Compared with NCS-1.0, the regular model brings an average BD rate change of {%,%,%,}, {%,%,%,}, and {-1.55%, -1.94%, -2.12%} for {Y, Cb, Cr} in RA, LB, and AI configurations. Compared with the NNVC-2.0 baseline, the regular model brings an average BD rate change of {%,%,%,}, {%,%,%,}, and {-8.68%, -21.49%, -22.09%} for {Y, Cb, Cr} in RA, LB, and AI configurations. Compact Model: The proposed model, with a regular size, includes 19 backbone blocks. Compared to NCS-1.0, the regular model results in an average BD rate change of {%,%,%,%}, {%,%,%,%}, and {%,%,%,%} for {Y, Cb, Cr} in RA, LB, and AI configurations. Compared to the NNVC-2.0 baseline, the regular model results in an average BD rate change of {%,%,%,%}, {%,%,%,%}, and {%,%,%,%} for {Y, Cb, Cr} in RA, LB, and AI configurations. Table 8. RA performance based on NCS-1.0 Table 9. LDB performance based on NCS-1.0 Table 10. AI performance based on NCS-1.0 Table 11. RA performance based on NNVC-2.0 Table 12. LDB performance based on NNVC-2.0 Table 13. AI performance based on NNVC-2.0 5.5. Conclusion This paper proposes a CNN-based loop filter network. The proposed method shows a favorable trade-off between codec performance and complexity. It is recommended to study the proposed method in EE.
[0092] Embodiments of the present disclosure relate to the use of NN models for encoding and decoding videos. One or more neural network (NN) models are trained as codec tools to improve the efficiency of video encoding and decoding. These NN-based codec tools can be used to replace or enhance the modules involved in the video codec. For example, the NN model can be used as an additional intra-frame prediction mode, transform kernel, or loop filter. The present invention details how to design the NN model by using external information such as prediction, partition, QP, etc. as attention.
[0093] It should be noted that the NN model can be used as any codec tool, such as NN-based intra / inter prediction, NN-based super-resolution, NN-based motion compensation, NN-based reference generation, NN-based fractional pixel interpolation, NN-based loop / post-processing filtering, etc.
[0094] In the present disclosure, the NN model can be any kind of NN architecture, such as a convolutional neural network or a fully connected neural network, or a combination of a convolutional neural network and a fully connected neural network.
[0095] In the following discussion, a video unit can be a sequence, a picture, a slice, a tile, a sub-picture, a codec tree unit (CTU), a codec tree block (CTB), a CTU row, a CTB row, one or more CUs / CBs, one or more CTUs / CTBs, one or more VPDUs (virtual pipeline data units), one or more codec units (CUs), one or more codec blocks (CBs), one or more CTUs, one or more CTBs, one or more virtual pipeline data units (VPDUs), a sub-region within a picture / slice / slice / tile, an inference block. A parent video unit represents a unit that is larger than a video unit. In some embodiments, a block can represent one or more samples, or one or more pixels. Typically, a parent unit will contain several video units. For example, when the video unit is a CTU, the parent unit can be a slice, a CTU row, multiple CTUs, etc.
[0096] The terms "frame" and "picture" may be used interchangeably. The terms "sample" and "pixel" may be used interchangeably.
[0097] Figure 24 A flowchart of a method 2400 for video processing according to an embodiment of the present disclosure is shown.
[0098] At block 2410, a neural network (NN) model for processing a video is obtained. The NN model includes at least one basic block. The basic block includes: a plurality of branches for processing the basic block's inputs in parallel; the branches include at least one convolutional layer and at least one activation layer; and a plurality of layers for serially processing a combination of the outputs of the plurality of branches, the plurality of layers including at least one convolutional layer and at least one activation layer. The plurality of branches are for processing the basic block's inputs in parallel; the branches include at least one convolutional layer and at least one activation layer; and a plurality of layers for serially processing a combination of the outputs of the plurality of branches, the plurality of layers including at least one convolutional layer and at least one activation layer.
[0099] At block 2420 , conversion between the current video block of the video and the bitstream of the video is performed according to the NN model.
[0100] Method 2400 enables the application of an efficient network architecture for video encoding and decoding that improves the performance-complexity trade-off. In this way, encoding and decoding performance can be further improved.
[0101] Figure 16A shows a basic block 1600A of a first type (type A) included in the NN model, and Figure 16B A basic block 1600B of the second type (type B) included in the NN model is shown. Figure 16A and Figure 16B In , regular rectangles represent convolutional layers (e.g., Layer1, Layer2, Layer5, Layer7), while rounded rectangles represent activation layers ((Layer3, Layer4, Layer6, Layer8). Figure 16A and Figure 16B In the figure, arrows show the data flow. According to the direction of the arrow, the output of the previous layer is used as the input of the next layer. out ,C in ,K hor ,K ver , S refer to the number of output channels, the number of input channels, the kernel size in the horizontal direction, the kernel size in the vertical direction, and the stride of the convolutional layer, respectively. Compared with (a), (b) includes a skip connection at the end, that is, the input of Lyaer1 (which is also the input of Lyaer2) is added to the output of Lyaer8. It should be understood that Figure 16A and Figure 16B The number of layers in is an example, and other variations are possible.
[0102] In some embodiments, within a basic block, a branch includes a single convolutional layer that receives the input of the basic block and a single activation layer that receives the output of the single convolutional layer. Examples of such basic blocks can be found in Figure 16A and Figure 16BThe single convolutional layer Lyaer1 in the branch is configured to receive the input of the basic block, and the single convolutional layer Lyaer2 in the other branch is also configured to receive the input of the basic block.
[0103] In some embodiments, within a basic block, the number of branches is two, an example of which is shown in Figure 16A and Figure 16B In some embodiments, within a basic block, multiple layers for serial processing include at least Figure 16A and / or Figure 16B The two convolutional layers in the basic block of , for example, Layer5 and Layer7.
[0104] In some embodiments, within a basic block, the multiple layers for serial processing also include Figure 16A and / or Figure 16B Two activation layers in the basic block, for example, Layer6 and Layer8.
[0105] In some embodiments, in (e.g., Figure 16A and / or Figure 16B In some embodiments, within a basic block, when the output of a previous layer is fed into multiple next layers, the input of each of the multiple next layers is the same as the output of the previous layer. In some embodiments, within a basic block, when the output of multiple previous layers is fed into the next layer, the input of the next layer is the concatenation of the outputs of the multiple previous layers along the channel dimension.
[0106] (For example, Figure 16A and / or Figure 16B The activation layers in the basic blocks can be configured in any flexible way.
[0107] In some embodiments, at least one of the activation layers included in the basic block is configured as a nonlinear function. In some embodiments, at least one of the activation layers included in the basic block is configured as a linear function.
[0108] In some embodiments, at least one activation layer included in the plurality of branches is configured as a parameterized rectified linear unit (PReLU). In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e., f1 and / or f2) in is PReLU.
[0109] In some embodiments, at least one activation layer included in the plurality of branches is configured as a leaky rectified linear unit (LReLU). In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e., f1 and / or f2) in is LReLU.
[0110] In some embodiments, at least one activation layer included in the plurality of branches is configured as a rectified linear unit (ReLU). In one example, Figure 16A and / or Figure 16B Layer3 and / or Layer4 (i.e., f1 and / or f2) in is ReLU.
[0111] In some embodiments, at least one activation layer included in the plurality of layers is configured as an identity mapping function, which means that the output and input of the layer are exactly the same. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is the identity mapping function (the output and input of this layer are exactly the same).
[0112] In some embodiments, at least one activation layer included in the plurality of layers is configured as a PReLU. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is PReLU.
[0113] In some embodiments, at least one activation layer included in the plurality of layers is configured as LReLU. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is LReLU.
[0114] In some embodiments, at least one activation layer included in the plurality of layers is configured as ReLU. In one example, Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is ReLU.
[0115] In one example, Figure 16A and / or Figure 16B All activation layers in are PReLU (parameterized rectified linear units). In one example, Figure 16A and / or Figure 16B All activation layers in are LReLU (leaky rectified linear units). In one example, Figure 16A and / or Figure 16B All activation layers in are ReLU (Rectified Linear Units).
[0116] In some embodiments, at least one activation layer included in the plurality of branches is configured as a nonlinear function, and at least one activation layer included in the plurality of layers is configured as a linear function. Figure 16Aand / or Figure 16B Lyaer3 and / or Lyaer4 (i.e. f1 and / or f2) in is a nonlinear function, and Figure 16A and / or Figure 16B Layer6 and / or Layer8 (ie, f3 and / or f4) are linear functions.
[0117] In some embodiments, at least one activation layer included in the plurality of branches is configured as a nonlinear function, and at least one activation layer included in the plurality of layers is configured as an identity mapping function. In one example, Figure 16A and / or Figure 16B v3 and / or Lyaer4 (i.e. f1 and / or f2) in are nonlinear functions, and Figure 16A and / or Figure 16B Layer6 and / or Layer8 (i.e., f3 and / or f4) in is an identity mapping function. In some embodiments, at least one activation layer included in the plurality of branches is configured as a parameterized rectified linear unit (PReLU), and at least one activation layer included in the plurality of layers is configured as an identity mapping function. In one example, Figure 16A and Figure 16B Layer3 and Layer4 (i.e. f1 and f2) in are PReLU (parameterized rectified linear units), and Figure 16A and Figure 16B Layer6 and Layer8 (i.e. f3 and f4) in are identity mapping functions.
[0118] In one example, which configuration to use with respect to activation layers in a basic block may be determined by decoding information. In one example, which configuration to use with respect to activation layers in a basic block may be determined by at least one syntax element signaled from an encoder to a decoder.
[0119] (For example, in Figure 16A and / or Figure 16B The convolutional layers in the basic blocks can be configured in any flexible way.
[0120] In some embodiments, the convolutional layers included in the basic block are configured with the same kernel size. In one example, the kernel size in all layers is the same, for example, 1×1, 3×3, 5×5, etc. In some embodiments, the convolutional layers included in the basic block are configured with different kernel sizes.
[0121] In some embodiments, convolutional layers included in multiple branches of a basic block are configured with different kernel sizes, and convolutional layers included in multiple layers of a basic block are configured with the same kernel size. In one example, the kernel sizes in Layer 1 and Layer 2 are different, while the kernel sizes in Layer 5 and Layer 7 are the same.
[0122] In some embodiments, if two branches are included in a basic block, each branch includes a convolution layer, and the multiple layers of the basic block include two convolution layers, the two convolution layers included in the two branches are configured with kernel sizes of 1×1 and 3×3, respectively, and the two convolution layers included in the multiple layers are both configured with a kernel size of 3×3. In one example, Figure 16A and / or Figure 16B The kernel sizes of Layer1 and Layer2 in are 1×1 and 3×3 respectively, while Figure 16A and / or Figure 16B The kernel sizes in Layer 5 and Layer 7 are both 3 × 3. In one example, the kernel sizes in Layer 1 and Layer 2 are 3 × 3 and 1 × 1 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3 × 3.
[0123] In some embodiments, the two convolutional layers included in the two branches are configured with kernel sizes of 1×1 and 5×5, respectively, and the two convolutional layers included in the multiple layers are both configured with kernel sizes of 3×3. Figure 16A and / or Figure 16B The kernel sizes of Layer1 and Layer2 in are 1×1 and 5×5 respectively, while Figure 16A and / or Figure 16B The kernel sizes in Layer 5 and Layer 7 are both 3 × 3. In one example, the kernel sizes in Layer 1 and Layer 2 are 5 × 5 and 1 × 1 respectively, while the kernel sizes in Layer 5 and Layer 7 are both 3 × 3.
[0124] In some embodiments, convolutional layers included in multiple branches of a basic block are configured with different kernel sizes, and convolutional layers included in multiple layers of a basic block are configured with different kernel sizes. In one example, the kernel sizes in Layer 1 and Layer 2 are different, and the kernel sizes in Layer 5 and Layer 7 are also different.
[0125] In some embodiments, if two branches are included in a basic block, each branch includes a convolution layer, and the multiple layers of the basic block include two convolution layers, the two convolution layers included in the two branches are configured with kernel sizes of 1×1 and 3×3, respectively, and the two convolution layers included in the multiple layers are configured with kernel sizes of 1×1 and 3×3, respectively. In one example, Figure 16A and / or Figure 16B The kernel sizes in Layer1 and Layer2 are 1×1 and 3×3 respectively, and Figure 16A and / or Figure 16B The kernel sizes in Layer5 and Layer7 are 1×1 and 3×3 respectively. In one example, Figure 16A and / or Figure 16B The kernel sizes in Layer 1 and Layer 2 are 3×3 and 1×1, respectively, and the kernel sizes in Layer 5 and Layer 7 are 1×1 and 3×3, respectively.
[0126] In some embodiments, the number of output channels in the convolutional layers included in the basic block is the same. In one example, the number of output channels (also referred to as output feature maps) in all layers is the same, for example, 32, 64, 96, 128, etc. In some embodiments, the number of output channels in the convolutional layers included in the basic block is different, for example, the number of output channels in all layers is different.
[0127] In some embodiments, the number of output channels of the convolution layers included in the multiple branches of the basic block is different, and the number of output channels of the convolution layers included in the multiple layers of the basic block is the same. In one example, Figure 16A and / or Figure 16B In an example, the number of output channels in Layer 1 and Layer 2 is different, while the number of output channels in Layer 5 and Layer 7 is the same. In one example, the number of output channels in Layer 1 is greater than the number of output channels in Layer 2. In another example, the number of output channels in Layer 1 is less than the number of output channels in Layer 2.
[0128] In some embodiments, the number of output channels of the convolutional layers included in the multiple branches of the basic block is different, and the number of output channels of the convolutional layers included in the multiple layers of the basic block is different. In one example, Figure 16A and / or Figure 16B The number of output channels in Layer1 and Layer2 is different, and Figure 16A and / or Figure 16BThe number of output channels in Layer 5 and Layer 7 is also different. In one example, the number of output channels in Layer 1 is greater than the number of output channels in Layer 2, and the number of output channels in Layer 5 is greater than the number of output channels in Layer 7. In one example, the number of output channels in Layer 1 is less than the number of output channels in Layer 2, and the number of output channels in Layer 5 is greater than the number of output channels in Layer 7.
[0129] In some embodiments, the convolutional layers included in the multiple branches of the basic block are configured with different kernel sizes and have different numbers of output channels; and wherein the convolutional layers included in the multiple layers of the basic block are configured with different kernel sizes and have the same number of output channels. In one example, the kernel sizes in Layer 1 and Layer 2 are different, Figure 16A and / or Figure 16B The kernel sizes in Layer5 and Layer7 are different. Figure 16A and / or Figure 16B The number of output channels in Layer1 and Layer2 is different. Figure 16A and / or Figure 16B The number of output channels in Layer5 and Layer7 is the same.
[0130] In some embodiments, two branches are included in a basic block, each branch includes a convolution layer, and multiple layers of the basic block include two convolution layers. The number of input channels to the two convolution layers included in the two branches of the basic block is denoted as N. The number of output channels in the two convolution layers included in the two branches is set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 and C2 is less than 1.0. The number of output channels of the two convolution layers included in the multiple layers is set to N, and the kernel sizes of the two convolution layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolution layers included in the multiple layers are set to 1×1 and 3×3, respectively. For example, the number of input channels of Layer1 (which is also the input channel of Layer2) is denoted as N. The number of output channels in Layer1 and Layer2 is set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 (e.g., 2.5) and C2 is less than 1.0 (e.g., 0.5). The number of output channels in Layer 5 and Layer 7 is set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively.
[0131] In some embodiments, two branches are included in a basic block, each branch includes a convolutional layer, a plurality of layers of the basic block include two convolutional layers, and the number of input channels to the two convolutional layers included in the two branches of the basic block is denoted as N. The number of output channels of the two convolutional layers included in the two branches is set to N×C1 and N×C2, respectively, where C1 is less than 1.0 and C2 is greater than 1.0. The number of output channels of the two convolutional layers included in the plurality of layers is set to N, and the kernel sizes of the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolutional layers included in the plurality of layers are set to 1×1 and 3×3, respectively. For example, the number of input channels of Layer 1 (which is also the input channel of Layer 2) is denoted as N. The number of output channels in Layer 1 and Layer 2 is set to N×C1 and N×C2, respectively, where C1 is less than 1.0 (e.g., 0.5) and C2 is greater than 1.0 (e.g., 2.5). The number of output channels in Layer 5 and Layer 7 is set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively.
[0132] In some embodiments, two branches are included in a basic block, each branch includes a convolutional layer, and multiple layers of the basic block include two convolutional layers, and multiple layers of the basic block include two convolutional layers. The number of output channels in the two convolutional layers included in the two branches is set to N×C1 and N×C2, respectively, where N is an integer, C1 is greater than 1.0 and C2 is less than 1.0. The number of output channels of the two convolutional layers included in the multiple layers is set to N, and the kernel sizes of the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolutional layers included in the multiple layers are set to 1×1 and 3×3, respectively. For example, given an integer N (e.g., 32, 64, 96, 128, etc.), the number of output channels in Layer 1 and Layer 2 is set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 (e.g., 2.5) and C2 is less than 1.0 (e.g., 0.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively.
[0133] In some embodiments, two branches are included in a basic block, each branch includes a convolution layer, and multiple layers of the basic block include two convolution layers, and multiple layers of the basic block include two convolution layers. The number of output channels in the two convolution layers included in the two branches is set to N×C1 and N×C2, respectively, where N is an integer, C1 is less than 1.0 and C2 is greater than 1.0. The number of output channels of the two convolution layers included in the multiple layers is set to N, and the kernel sizes of the two convolution layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolution layers included in the multiple layers are set to 1×1 and 3×3, respectively. For example, given an integer N (e.g., 32, 64, 96, 128, etc.), the number of output channels in Layer 1 and Layer 2 is set to N×C1 and N×C2, respectively, where C1 is less than 1.0 (e.g., 0.5) and C2 is greater than 1.0 (e.g., 2.5). The number of output channels in Layer 5 and Layer 7 is both set to N. The kernel sizes in Layer 1 and Layer 2 are 1×1 and 3×3, respectively, and the kernel sizes in Layer 5 and Layer 7 are set to 1×1 and 3×3, respectively.
[0134] In some embodiments, a configuration associated with multiple branches in a basic block, a configuration associated with multiple layers in a basic block, a configuration associated with an activation layer in a basic block, and / or a configuration associated with a convolutional layer in a basic block is determined based on at least one of: decoding information of the video, or at least one syntax element transmitted by a signal from an encoder of the video to a decoder of the video.
[0135] In some embodiments, the above configurations related to activation layers and convolutional layers can be combined. Specifically, the activation layer can be configured in any of the above embodiments, while the convolutional layer can be configured using any of the above embodiments.
[0136] Figures 16A-16B The activation layers and convolutional layers shown can be jointly configured in certain ways to achieve better performance.
[0137] In some embodiments, within a basic block, activation layers included in multiple branches of the basic block are configured as nonlinear functions, and activation layers included in multiple layers of the basic block are configured as identity mapping functions; and convolution layers included in multiple branches of the basic block are configured with different kernel sizes, and convolution layers included in multiple layers of the basic block are configured with different kernel sizes; and the number of output channels of the convolution layers included in multiple branches of the basic block is different, and the number of output channels of the convolution layers included in multiple layers of the basic block is the same. In one example, Figure 16A and / or Figure 16BLayer3 and Layer4 (i.e. f1 and f2) in are nonlinear functions, and Figure 16A and / or Figure 16B Layer6 and Layer8 (i.e. f3 and f4) in are the same mapping functions. The kernel sizes in Layer1 and Layer2 are different. Figure 16A and / or Figure 16B The kernel sizes in Layer5 and Layer7 are different, the number of output channels in Layer1 and Layer2 are different, and the number of output channels in Layer5 and Layer7 are the same.
[0138] In some embodiments, the activation layers included in the plurality of branches of the basic block are configured as at least one of the following: a parameterized rectified linear unit (PReLU), a leaky rectified linear unit (LReLU), or a rectified linear unit (ReLU). In one example, Figure 16A and Figure 16B Layer3 and Layer4 (i.e., f1 and f2) in are PReLU (parameterized rectified linear units). In one example, Figure 16A and Figure 16B Layer3 and Layer4 (i.e., f1 and f2) in are LReLU (leaky rectified linear units). In one example, Figure 16A and Figure 16B Layer3 and Layer4 (i.e. f1 and f2) in are ReLU (rectified linear units).
[0139] In some embodiments, among multiple branches of a basic block, the kernel size of a first convolutional layer included in a first branch is smaller than the kernel size of a second convolutional layer included in a second branch, and a first number of output channels in the first convolutional layer is greater than a second number of output channels in the second convolutional layer, and among multiple layers of the basic block, the kernel size of a preceding convolutional layer is smaller than the kernel size of a succeeding convolutional layer. In one example, the kernel size in Layer 1 is smaller than the kernel size in Layer 2, and the kernel size in Layer 5 is smaller than the kernel size in Layer 7. The number of output channels in Layer 1 is greater than the number of output channels in Layer 2.
[0140] In some embodiments, among multiple branches of a basic block, a kernel size in a first convolution layer included in a first branch is larger than a kernel size in a second convolution layer included in a second branch, and a first number of output channels in the first convolution layer is smaller than a second number of output channels in the second convolution layer, and among multiple layers of the basic block, a kernel size in a previous convolution layer is smaller than a kernel size in a subsequent convolution layer. In one example, the kernel size is larger than the kernel size, wherein the kernel size is smaller than the number of output channels.
[0141] In some embodiments, at least one basic block included in the NN model includes at least one basic block of a first type and / or at least one basic block of a second type, wherein the basic block of the first type does not have a skip connection that adds the input of the basic block to the output of the last layer of the basic block; and wherein the basic block of the second type has a skip connection that adds the input of the basic block to the output of the last layer of the basic block. For example, the NN model may include Figures 16A-16B In one example, the NN model includes one or more basic blocks shown in FIG. Figure 16A In one example, the NN model includes Figure 16B In one example, the NN model includes Figure 16A At least one basic block shown and Figure 16B At least one basic block shown in .
[0142] In some embodiments, the NN model includes a head part, a backbone part, and a tail part, wherein the head part is configured to extract features from the input of the NN model, the backbone part is configured for further feature mapping, and the tail part is configured to transform the output features of the backbone part into the output of the NN model. Figure 17 As shown in the architecture 1700, the NN model may include three parts, wherein the head part is responsible for extracting features from the input of the NN model, which are then fed into the backbone part for further feature mapping. The tail part transforms the output features of the backbone part into the final output.
[0143] In some embodiments, at least one of the head portion, the backbone portion, or the tail portion each includes a first number of basic blocks of a first type connected in series. Figure 18 The example shown, Figure 18 A stack comprising basic blocks 1800, wherein basic block type A is Figure 16A The basic blocks shown, and M a is the number of stacked blocks.
[0144] In some embodiments, at least one of the head portion, the backbone portion, or the tail portion each includes a second number of basic blocks of the second type connected in series. Figure 19 The example shown, Figure 19 A stack comprising basic blocks 1900, wherein basic block type B is Figure 16B The basic blocks shown, and M b is the number of blocks to be stacked. M b Can be set to any suitable number.
[0145] In some embodiments, at least one of the head portion, the backbone portion, or the tail portion each includes a first number of basic blocks of a first type connected in series, followed by a second number of basic blocks of a second type connected in series. Figure 20 The example shown, Figure 20 A stack comprising basic blocks 2000, wherein basic block type A is Figure 16A The basic block shown, and the basic block type B is Figure 16B The basic blocks shown. M a and M b is the number of blocks to be stacked. M a and M b Can be set to any suitable number.
[0146] In some embodiments, at least one of the head portion, the backbone portion, or the tail portion each includes a second number of basic blocks of the second type connected in series, followed by a first number of basic blocks of the first type connected in series. Figure 21 The example shown, Figure 21 A stack comprising basic blocks 2100, wherein basic block type B is Figure 16B The basic block shown, and the basic block type A is Figure 16A The basic blocks shown. M b and M a is the number of blocks to be stacked. M b and M a Can be set to any suitable number.
[0147] In one example, the header section begins with Figure 18 In one example, the head portion is designed in the manner shown. Figure 19 In one example, the head portion is designed in the manner shown. Figure 20 In one example, the head portion is designed in the manner shown. Figure 21 The manner shown is designed.
[0148] Similarly, in one example, the backbone portion is Figure 18 In one example, the backbone portion is designed in the manner shown. Figure 19 In one example, the backbone portion is designed in the manner shown. Figure 20 In one example, the backbone portion is designed in the manner shown. Figure 21 The method shown is designed. The number of basic blocks in the backbone part M a and / or M b The number of basic blocks M in the header or tail part can be a and / or M b Same or different.
[0149] Similarly, in one example, the tail portion begins with Figure 18 In one example, the tail portion is designed in the manner shown. Figure 19 In one example, the tail portion is designed in the manner shown. Figure 20 In one example, the tail portion is designed in the manner shown. Figure 21 The way shown is designed. The number of basic blocks in the tail part M a and / or M b It can be combined with the number of basic blocks M in the header part or the backbone part a and / or M b Same or different.
[0150] In some embodiments, the head portion includes a second number of basic blocks of the second type connected in series, the backbone portion includes a second number of basic blocks of the second type connected in series, followed by a first number of basic blocks of the first type connected in series, and the tail portion includes a conventional convolutional layer. Figure 18 The backbone is designed in the manner shown. Figure 21 The method shown is designed, and the tail is a regular convolutional layer.
[0151] In some embodiments, integer operations are applied in the NN model. For example, only integer operations are applied in the proposed architecture. In some embodiments, floating-point operations are not applied in the NN model. In some embodiments, division operations are not applied in the NN model.
[0152] In some embodiments, the integer operations applied in the NN model include at least one of the following: an addition operation, a multiplication operation, a shift operation, a rounding operation, and a clipping operation.
[0153] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, which is generated by a method executed by a video processing device. The method includes: obtaining a neural network (NN) model for processing a video, the NN model including at least one basic block, wherein the basic block includes: multiple branches for processing the input of the basic block in parallel, the branches including at least one convolutional layer and at least one activation layer, and multiple layers for serially processing a combination of outputs of the multiple branches, the multiple layers including at least one convolutional layer and at least one activation layer; and generating a bitstream of the video according to the NN model.
[0154] According to another embodiment of the present disclosure, a method for storing a video bitstream is provided. The method includes: obtaining a neural network (NN) model for processing a video, the NN model including at least one basic block, wherein the basic block includes: multiple branches for processing the basic block's input in parallel, the branches including at least one convolutional layer and at least one activation layer, and multiple layers for serially processing a combination of the outputs of the multiple branches, the multiple layers including at least one convolutional layer and at least one activation layer; generating a video bitstream according to the NN model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0155] The implementation of the present disclosure can be described with reference to the following items, and the features thereof can be combined in any reasonable manner.
[0156] Item 1. A method for video processing, comprising: obtaining a neural network (NN) model for processing a video, the NN model comprising at least one basic block, wherein the basic block comprises: a plurality of branches for processing an input of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for processing a combination of outputs of the plurality of branches in serial, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and performing conversion between a current video block of the video and a bitstream of the video according to the NN model.
[0157] Item 2. The method of Item 1, wherein within a basic block, a branch comprises a single convolutional layer receiving the input of the basic block and a single activation layer receiving the output of the single convolutional layer.
[0158] Item 3. A method according to Item 1 or 2, wherein within a basic block, the number of branches is 2; and / or wherein within a basic block, the multiple layers for serial processing include at least two convolutional layers; and / or wherein within a basic block, the multiple layers for serial processing also include two activation layers.
[0159] Item 4. A method according to any one of Items 1 to 3, wherein within a basic block, where the output of a previous layer is fed into a plurality of next layers, the input of each of the plurality of next layers is the same as the output of the previous layer.
[0160] Item 5. A method according to any one of Items 1 to 4, wherein within a basic block, where the outputs of multiple previous layers are fed into a next layer, the input of the next layer is the concatenation of the outputs of the multiple previous layers along the channel dimension.
[0161] Item 6. A method according to any one of Items 1 to 5, wherein at least one activation layer included in the basic block is configured as at least one of the following: a nonlinear function or a linear function; and / or wherein at least one activation layer included in the multiple branches is configured as at least one of the following: a parameterized rectified linear unit (PReLU), a leaky rectified linear unit (LReLU) or a rectified linear unit (ReLU); and / or wherein at least one activation layer included in the multiple layers is configured as at least one of the following: an identity mapping function, a PReLU, a LReLU or a ReLU.
[0162] Item 7. A method according to any one of Items 1 to 6, wherein at least one activation layer included in the multiple branches is configured as a nonlinear function, and at least one activation layer included in the multiple layers is configured as a linear function; and / or wherein at least one activation layer included in the multiple branches is configured as a nonlinear function, and at least one activation layer included in the multiple layers is configured as an identity mapping function; and / or wherein at least one activation layer included in the multiple branches is configured as a parameterized rectified linear unit (PReLU), and at least one activation layer included in the multiple layers is configured as an identity mapping function.
[0163] Item 8. A method according to any one of Items 1 to 7, wherein the convolution layers included in the basic block are configured with the same kernel size, or wherein the convolution layers included in the basic block are configured with different kernel sizes, or wherein the convolution layers included in the multiple branches of the basic block are configured with different kernel sizes, and the convolution layers included in the multiple layers of the basic block are configured with the same kernel size, or wherein the convolution layers included in the multiple branches of the basic block are configured with different kernel sizes, and the convolution layers included in the multiple layers of the basic block are configured with different kernel sizes.
[0164] Item 9. The method according to Item 8, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the multiple layers of the basic block include two convolutional layers; and wherein the two convolutional layers included in the two branches are respectively configured with kernel sizes of 1×1 and 3×3, and the two convolutional layers included in the multiple layers are both configured with a kernel size of 3×3; or wherein the two convolutional layers included in the two branches are respectively configured with kernel sizes of 1×1 and 5×5, and the two convolutional layers included in the multiple layers are both configured with a kernel size of 3×3.
[0165] Item 10. A method according to Item 8, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the multiple layers of the basic block include two convolutional layers; and wherein the two convolutional layers included in the two branches are respectively configured with kernel sizes of 1×1 and 3×3, and the two convolutional layers included in the multiple layers are respectively configured with kernel sizes of 1×1 and 3×3.
[0166] Item 11. A method according to any one of Items 1 to 10, wherein the number of output channels of the convolutional layers included in the basic block is the same; or wherein the number of output channels of the convolutional layers included in the basic block is different; or wherein the number of output channels of the convolutional layers included in the multiple branches of the basic block is different, and the number of output channels of the convolutional layers included in the multiple layers of the basic block is the same; or wherein the number of output channels of the convolutional layers included in the multiple branches of the basic block is different, and the number of output channels of the convolutional layers included in the multiple layers of the basic block is different.
[0167] Item 12. A method according to any one of Items 1 to 11, wherein the convolutional layers included in the multiple branches of the basic block are configured with different kernel sizes and are configured with different numbers of output channels; and wherein the convolutional layers included in the multiple layers of the basic block are configured with different kernel sizes and are configured with the same number of output channels.
[0168] Item 13. A method according to Item 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, the multiple layers of the basic block include two convolutional layers, and the number of input channels of the two convolutional layers included in the two branches of the basic block is denoted as N, the number of output channels of the two convolutional layers included in the two branches is set to N×C1 and N×C2, respectively, where C1 is greater than 1.0 and C2 is less than 1.0, the number of output channels of the two convolutional layers included in the multiple layers is set to N, and the kernel sizes of the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolutional layers included in the multiple layers are set to 1×1 and 3×3, respectively.
[0169] Item 14. A method according to Item 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, the multiple layers of the basic block include two convolutional layers, and the number of input channels of the two convolutional layers included in the two branches of the basic block is denoted as N, the number of output channels of the two convolutional layers included in the two branches is set to N×C1 and N×C2, respectively, where C1 is less than 1.0 and C2 is greater than 1.0, the number of output channels of the two convolutional layers included in the multiple layers is set to N, and the kernel sizes of the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes of the two convolutional layers included in the multiple layers are set to 1×1 and 3×3, respectively.
[0170] Item 15. A method according to Item 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the multiple layers of the basic block include two convolutional layers, the number of output channels in the two convolutional layers included in the two branches are set to N×C1 and N×C2, respectively, where N is an integer, C1 is greater than 1.0 and C2 is less than 1.0; and the number of output channels in the two convolutional layers included in the multiple layers is both set to N; and the kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the multiple layers are set to 1×1 and 3×3, respectively.
[0171] Item 16. A method according to Item 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the multiple layers of the basic block include two convolutional layers, the number of output channels in the two convolutional layers included in the two branches is set to N×C1 and N×C2, respectively, where N is an integer, C1 is less than 1.0 and C2 is greater than 1.0; the number of output channels in the two convolutional layers included in the multiple layers is set to N; and the kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the multiple layers are set to 1×1 and 3×3, respectively.
[0172] Item 17. A method according to any one of Items 1 to 16, wherein the configuration associated with the multiple branches in the basic block, the configuration associated with the multiple layers in the basic block, the configuration associated with the activation layer in the basic block and / or the configuration associated with the convolutional layer in the basic block is determined based on at least one of the following: decoded information of the video, or at least one syntax element transmitted from the encoder of the video to the decoder of the video via a signal.
[0173] Item 18. A method according to any one of Items 1 to 17, wherein, within a basic block, the activation layers included in the multiple branches of the basic block are configured as nonlinear functions, and the activation layers included in the multiple layers of the basic block are configured as identity mapping functions; and the convolution layers included in the multiple branches of the basic block are configured with different kernel sizes, and the convolution layers included in the multiple layers of the basic block are configured with different kernel sizes; and the number of output channels of the convolution layers included in the multiple branches of the basic block is different, and the number of output channels of the convolution layers included in the multiple layers of the basic block is the same.
[0174] Item 19. The method according to Item 18, wherein the activation layers included in the multiple branches of the basic block are configured as at least one of the following: a parameterized rectified linear unit (PReLU), a leaky rectified linear unit (LReLU), or a rectified linear unit (ReLU); and / or wherein, in the multiple branches of the basic block, a kernel size in a first convolutional layer included in a first branch is smaller than a kernel size in a second convolutional layer included in a second branch, and a first number of output channels in the first convolutional layer is larger than a second number of output channels in the second convolutional layer, and in the multiple layers of the basic block, a kernel size in a previous convolutional layer is smaller than a kernel size in a subsequent convolutional layer; or, wherein, in the multiple branches of the basic block, a kernel size in a first convolutional layer included in the first branch is larger than a kernel size in a second convolutional layer included in the second branch, and the first number of output channels in the first convolutional layer is smaller than the second number of output channels in the second convolutional layer, and in the multiple layers of the basic block, a kernel size in a previous convolutional layer is smaller than a kernel size in a subsequent convolutional layer.
[0175] Item 20. A method according to any one of Items 1 to 19, wherein the at least one basic block included in the NN model includes at least one basic block of a first type and / or at least one basic block of a second type; wherein the basic block of the first type does not have a jump connection that adds the input of the basic block to the output of the last layer of the basic block; and wherein the basic block of the second type has a jump connection that adds the input of the basic block to the output of the last layer of the basic block.
[0176] Item 21. A method according to any one of Items 1 to 20, wherein the NN model includes a head part, a backbone part and a tail part, wherein the head part is configured to extract features from the input of the NN model, the backbone part is configured for further feature mapping, and the tail part is configured to transform the output features of the backbone part into the output of the NN model.
[0177] Item 22. A method according to item 21, wherein at least one of the head part, the backbone part or the tail part each includes a first number of basic blocks of the first type connected in series; or wherein at least one of the head part, the backbone part or the tail part each includes a second number of basic blocks of the second type connected in series; or wherein at least one of the head part, the backbone part or the tail part each includes a first number of basic blocks of the first type connected in series, followed by a second number of basic blocks of the second type connected in series; or wherein at least one of the head part, the backbone part or the tail part each includes a second number of basic blocks of the second type connected in series, followed by a first number of basic blocks of the first type connected in series, or wherein the head part includes a second number of basic blocks of the second type connected in series, the backbone part includes the second number of basic blocks of the second type connected in series, followed by a first number of basic blocks of the first type connected in series, and the tail part includes a conventional convolutional layer.
[0178] Item 23. A method according to any one of Items 1 to 22, wherein integer operations are applied in the NN model; and wherein floating-point operations are not applied in the NN model; and wherein division operations are not applied in the NN model; and / or wherein the integer operations applied in the NN model include at least one of the following: addition operations, multiplication operations, shift operations, rounding operations, and clipping operations.
[0179] Item 24. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 23.
[0180] Item 25. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 23.
[0181] Item 26. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream being generated by a method executed by a video processing device, wherein the method comprises: obtaining a neural network (NN) model for processing a video, the NN model comprising at least one basic block, wherein the basic block comprises: a plurality of branches for processing inputs of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for processing a combination of outputs of the plurality of branches in serial, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and generating a bitstream of the video according to the NN model.
[0182] Item 27. A method for storing a bitstream of a video, comprising: obtaining a neural network (NN) model for processing a video, the NN model comprising at least one basic block, wherein the basic block comprises: a plurality of branches for processing an input of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for processing a combination of outputs of the plurality of branches in serial, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and generating a bitstream of the video based on the NN model; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0183] Figure 25 A block diagram of a computing device 2500 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2500 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0184] It should be understood that Figure 25 The computing device 2500 shown in FIG. 2 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0185] like Figure 25As shown, computing device 2500 comprises a general computing device 2500. Computing device 2500 may include at least one or more processors or processing units 2510, memory 2520, storage unit 2530, one or more communication units 2540, one or more input devices 2550, and one or more output devices 2560.
[0186] In some embodiments, computing device 2500 can be implemented as any user terminal or server terminal with computing capability. A server terminal can be a server, a large computing device, etc. provided by a service provider. A user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 2500 can support any type of interface to a user (such as a "wearable" circuit device, etc.).
[0187] The processing unit 2510 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 2520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 2500. The processing unit 2510 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0188] The computing device 2500 typically includes various computer storage media. Such media can be any media accessible by the computing device 2500, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 2520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 2530 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 2500.
[0189] The computing device 2500 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 25 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0190] The communication unit 2540 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 2500 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 2500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0191] Input device 2550 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 2560 may be one or more of various output devices, such as a display, speaker, printer, and the like. Computing device 2500 may also communicate with one or more external devices (not shown) via communication unit 2540, such as storage devices and display devices, one or more devices that enable a user to interact with computing device 2500, or any device that enables computing device 2500 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like), if desired. Such communication may occur via an input / output (I / O) interface (not shown).
[0192] In some embodiments, some or all components of the computing device 2500 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0193] In an embodiment of the present disclosure, the computing device 2500 may be used to implement video encoding / decoding. The memory 2520 may include one or more video encoding / decoding modules 2525 having one or more program instructions. These modules are accessible and executable by the processing unit 2510 to perform the functions of the various embodiments described herein.
[0194] In an example embodiment performing video encoding, an input device 2550 may receive video data as input 2570 to be encoded. The video data may be processed, for example, by a video codec module 2525 to generate an encoded bitstream. The encoded bitstream may be provided as output 2580 via an output device 2560.
[0195] In an example embodiment performing video decoding, an input device 2550 may receive an encoded bitstream as input 2570. The encoded bitstream may be processed, for example, by a video codec module 2525 to generate decoded video data. The decoded video data may be provided as output 2580 via an output device 2560.
[0196] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, include: Obtain a neural network (NN) model for processing a video, wherein the NN model includes at least one basic block, wherein the basic block includes: A plurality of branches for processing the input of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for serially processing a combination of outputs of the plurality of branches, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and According to the NN model, a conversion between a current video block of the video and a bitstream of the video is performed.
2. The method of claim 1, wherein within a basic block, a branch comprises a single convolutional layer receiving the input of the basic block and a single activation layer receiving the output of the single convolutional layer.
3. The method according to claim 1 or 2, wherein within a basic block, the number of branches is 2; and / or wherein within the basic block, the plurality of layers for serial processing include at least two convolutional layers; and / or Wherein within the basic block, the plurality of layers for serial processing further include two activation layers.
4. The method according to any one of claims 1 to 3, wherein within a basic block, when the output of a previous layer is fed into a plurality of next layers, the input of each of the plurality of next layers is the same as the output of the previous layer.
5. The method according to any one of claims 1 to 4, wherein within a basic block, when the outputs of multiple previous layers are fed into a next layer, the input of the next layer is the concatenation of the outputs of the multiple previous layers along the channel dimension.
6. The method according to any one of claims 1 to 5, wherein at least one activation layer included in the basic block is configured as at least one of the following: a nonlinear function or a linear function; and / or At least one activation layer included in the plurality of branches is configured as at least one of the following: a parameterized rectified linear unit (PReLU), a leaky rectified linear unit (LReLU) or a rectified linear unit (ReLU); and / or At least one activation layer included in the multiple layers is configured as at least one of the following: an identity mapping function, PReLU, LReLU or ReLU.
7. The method according to any one of claims 1 to 6, wherein at least one activation layer included in the plurality of branches is configured as a nonlinear function, and at least one activation layer included in the plurality of layers is configured as a linear function. ; and / or wherein at least one activation layer included in the plurality of branches is configured as a nonlinear function, and at least one activation layer included in the plurality of layers is configured as an identity mapping function; and / or At least one activation layer included in the multiple branches is configured as a parameterized rectified linear unit (PReLU), and at least one activation layer included in the multiple layers is configured as an identity mapping function.
8. The method according to any one of claims 1 to 7, wherein the convolutional layers included in the basic block are configured with the same kernel size, or where the convolutional layers included in the basic block are configured with different kernel sizes, or wherein the convolutional layers included in the plurality of branches of the basic block are configured with different kernel sizes, and the convolutional layers included in the plurality of layers of the basic block are configured with the same kernel size, or The convolutional layers included in the multiple branches of the basic block are configured with different kernel sizes, and the convolutional layers included in the multiple layers of the basic block are configured with different kernel sizes.
9. The method of claim 8, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the plurality of layers of the basic block includes two convolutional layers; and wherein the two convolutional layers included in the two branches are configured with kernel sizes of 1×1 and 3×3, respectively, and the two convolutional layers included in the plurality of layers are both configured with a kernel size of 3×3; or The two convolutional layers included in the two branches are configured with kernel sizes of 1×1 and 5×5, respectively, and the two convolutional layers included in the multiple layers are both configured with a kernel size of 3×3.
10. The method of claim 8, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the plurality of layers of the basic block includes two convolutional layers; and The two convolutional layers included in the two branches are configured with kernel sizes of 1×1 and 3×3, respectively, and the two convolutional layers included in the multiple layers are configured with kernel sizes of 1×1 and 3×3, respectively.
11. The method according to any one of claims 1 to 10, wherein the number of output channels in the convolutional layers included in the basic block is the same; or wherein the number of output channels in the convolutional layers included in the basic block is different; or wherein the numbers of output channels in the convolutional layers included in the plurality of branches of the basic block are different, and the numbers of output channels in the convolutional layers included in the plurality of layers of the basic block are the same; or The number of output channels of the convolutional layers included in the multiple branches of the basic block is different, and the number of output channels of the convolutional layers included in the multiple layers of the basic block is different.
12. The method according to any one of claims 1 to 11, wherein the convolutional layers included in the plurality of branches of the basic block are configured with different kernel sizes and are configured with different numbers of output channels; and The convolutional layers included in the plurality of layers of the basic block are configured with different kernel sizes and are configured with the same number of output channels.
13. The method according to claim 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, the multiple layers of the basic block include two convolutional layers, and the number of input channels of the two convolutional layers included in the two branches of the basic block is recorded as N, The number of output channels in the two convolutional layers included in the two branches is set to N×C 1 and N×C 2 , where C 1 Greater than 1.0 and C 2 Less than 1.0, The number of output channels in the two convolutional layers included in the plurality of layers is set to N, and The kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the plurality of layers are set to 1×1 and 3×3, respectively.
14. The method according to claim 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, the multiple layers of the basic block include two convolutional layers, and the number of input channels of the two convolutional layers included in the two branches of the basic block is recorded as N, The number of output channels in the two convolutional layers included in the two branches is set to N×C 1 and N×C 2 , where C 1 Less than 1.0 and C 2 greater than 1.0, The number of output channels in the two convolutional layers included in the plurality of layers is set to N, and The kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the plurality of layers are set to 1×1 and 3×3, respectively.
15. The method of claim 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the plurality of layers of the basic block includes two convolutional layers, The number of output channels in the two convolutional layers included in the two branches is set to N×C 1 and N×C 2 , where N is an integer, C 1 Greater than 1.0 and C 2 less than 1.0; and The numbers of output channels in the two convolutional layers included in the plurality of layers are both set to N; and The kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the plurality of layers are set to 1×1 and 3×3, respectively.
16. The method of claim 12, wherein two branches are included in a basic block, each of the branches includes a convolutional layer, and the plurality of layers of the basic block includes two convolutional layers, The number of output channels in the two convolutional layers included in the two branches is set to N×C 1 and N×C 2 , where N is an integer, C 1 Less than 1.0 and C 2 Greater than 1.0; The number of output channels in the two convolutional layers included in the plurality of layers is set to N; and The kernel sizes in the two convolutional layers included in the two branches are 1×1 and 3×3, respectively, and the kernel sizes in the two convolutional layers included in the plurality of layers are set to 1×1 and 3×3, respectively.
17. The method according to any one of claims 1 to 16, wherein the configuration associated with the plurality of branches in a basic block, the configuration associated with the plurality of layers in a basic block, the configuration associated with the activation layer in a basic block, and / or the configuration associated with the convolutional layer in a basic block is determined based on at least one of the following: decoded information of the video, or At least one syntax element is signaled from an encoder of the video to a decoder of the video.
18. The method according to any one of claims 1 to 17, wherein within a basic block, The activation layers included in the plurality of branches of the basic block are configured as nonlinear functions, and the activation layers included in the plurality of layers of the basic block are configured as identity mapping functions; and The convolutional layers included in the plurality of branches of the basic block are configured with different kernel sizes, and the convolutional layers included in the plurality of layers of the basic block are configured with different kernel sizes; and The numbers of output channels in the convolutional layers included in the plurality of branches of the basic block are different, and the numbers of output channels in the convolutional layers included in the plurality of layers of the basic block are the same.
19. The method of claim 18, wherein the activation layers included in the plurality of branches of the basic block are configured as at least one of: a parameterized rectified linear unit (PReLU), a leaky rectified linear unit (LReLU), or a rectified linear unit (ReLU); and / or wherein, among the multiple branches of the basic block, a kernel size in a first convolutional layer included in a first branch is smaller than a kernel size in a second convolutional layer included in a second branch, and a first number of output channels in the first convolutional layer is larger than a second number of output channels in the second convolutional layer, and among the multiple layers of the basic block, a kernel size in a preceding convolutional layer is smaller than a kernel size in a succeeding convolutional layer; or, Among the multiple branches of the basic block, the kernel size of the first convolution layer included in the first branch is larger than the kernel size of the second convolution layer included in the second branch, and the first number of output channels in the first convolution layer is smaller than the second number of output channels in the second convolution layer, and among the multiple layers of the basic block, the kernel size in the preceding convolution layer is smaller than the kernel size in the succeeding convolution layer.
20. The method according to any one of claims 1 to 19, wherein the at least one basic block included in the NN model comprises at least one basic block of a first type and / or at least one basic block of a second type; wherein the first type of basic block does not have a skip connection that adds the input of the basic block to the output of the last layer of the basic block; and The second type of basic block has a skip connection that adds the input of the basic block to the output of the last layer of the basic block.
21. The method according to any one of claims 1 to 20, wherein the NN model comprises a head part, a backbone part and a tail part, The head part is configured to extract features from the input of the NN model, the backbone part is configured for further feature mapping, and the tail part is configured to transform the output features of the backbone part into the output of the NN model.
22. The method of claim 21, wherein at least one of the head portion, the backbone portion, or the tail portion each comprises a first number of basic blocks of the first type connected in series; or wherein at least one of the head portion, the backbone portion or the tail portion each comprises a second number of basic blocks of the second type connected in series; or wherein at least one of the head portion, the backbone portion or the tail portion each comprises a first number of basic blocks of the first type connected in series followed by a second number of basic blocks of the second type connected in series; or wherein at least one of the head portion, the backbone portion or the tail portion each comprises a second number of basic blocks of the second type connected in series followed by a first number of basic blocks of the first type connected in series, or wherein the head part comprises a second number of basic blocks of the second type connected in series, the backbone part comprises the second number of basic blocks of the second type connected in series, followed by a first number of basic blocks of the first type connected in series, and the tail part comprises a conventional convolutional layer.
23. The method according to any one of claims 1 to 22, wherein integer operations are applied in the NN model; and wherein floating point operations are not applied in the NN model; and wherein a division operation is not applied in the NN model; and / or The integer operation applied in the NN model includes at least one of the following: addition operation, multiplication operation, shift operation, rounding operation, and limiting operation.
24. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 23.
25. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 23.
26. A non-transitory computer-readable recording medium storing a bit stream of a video, the bit stream being generated by a method executed by a video processing device, wherein the method include: Obtain a neural network (NN) model for processing a video, wherein the NN model includes at least one basic block, wherein the basic block includes: A plurality of branches for processing the input of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for serially processing a combination of outputs of the plurality of branches, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and According to the NN model, a bit stream of the video is generated.
27. A method for storing a bit stream of a video, include: Obtain a neural network (NN) model for processing a video, wherein the NN model includes at least one basic block, wherein the basic block includes: A plurality of branches for processing the input of the basic block in parallel, the branches comprising at least one convolutional layer and at least one activation layer, and a plurality of layers for serially processing a combination of outputs of the plurality of branches, the plurality of layers comprising at least one convolutional layer and at least one activation layer; and Generating a bitstream of the video according to the NN model; and The bit stream is stored in a non-transitory computer-readable recording medium.