Method and device for video processing and medium
By introducing a neural network-based loop filter into the video encoding and decoding process, the filtering performance of video processing is optimized and the filter complexity is reduced, thus addressing the need for improved encoding and decoding efficiency in existing technologies.
Patent Information
- Application Number
- CN202480027761.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-23
- Filing Date
- 2024-04-22
- Publication Date
- 2025-12-12
AI Technical Summary
There is room for improvement in the efficiency of existing video encoding and decoding technologies, especially when processing video data, where the complexity and performance of existing filters need to be optimized.
By employing a neural network-based loop filter, different convolution types, multi-scaling neural network structures, transformer-based structures, and combinations of non-neural network filters and neural network filters are used in the conversion process between video units and bitstreams to improve filtering performance and reduce filter complexity.
It improves the filtering performance of video processing, while reducing the complexity of the filter and improving the overall efficiency of video encoding and decoding.
Smart Images

Figure CN121128182A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to neural network-based loop filtering for video encoding and decoding design. Background Technology
[0002] Digital video capabilities are now being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: during the conversion between video units and bitstreams of video units, determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs to the neural network filter; convolutions with kernel sizes are decomposed into combinations of multiple convolutions with smaller kernel sizes; which side information is to be used as input to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with a neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to the video units; and performing the conversion based on the filtered video units. In this way, the filtering performance can be improved and the filter complexity reduced.
[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs to the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information will be used as inputs to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with a neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; and generating a bitstream based on the filtered video units.
[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs of the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information is to be used as input to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with a neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; generating a bitstream based on the filtered video units; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 An example of raster scan strip segmentation of an image is shown; Figure 5 An example of rectangular strip segmentation of an image is shown; Figure 6 Examples of images segmented into pieces, bricks, and rectangular strips are shown; Figure 7A An example image showing CTB spanning the bottom image boundary is shown; Figure 7B An example image showing CTB crossing the right edge of the image is shown; Figure 7C An example image showing CTB crossing the bottom right edge of the image is shown; Figure 8 An example of a VVC encoder block diagram is shown; Figure 9 The image samples and horizontal and vertical block boundaries on an 8×8 grid are shown, as well as the non-overlapping blocks of the 8×8 samples; Figure 10 The pixels involved in the filter on / off decision and strong / weak filter selection are shown; Figures 11A to 11D Example diagrams illustrating four 1-D orientation patterns used for EO sample point classification are shown; Figures 12A to 12C An example diagram showing an example of the shape of a GALF filter is provided; Figures 13A to 13C An example diagram is shown illustrating an example of a relative coordinator supported by a 5×5 diamond filter; Figure 14 An example diagram is shown illustrating an example of the relative coordinates supported by a 5×5 diamond filter; Figure 15A An example diagram illustrating the architecture of the proposed CNN filter is shown; Figure 15B An example diagram illustrating the construction of a ResBlock in a CNN filter is shown; Figure 16 An example neural network with two branches is shown; Figure 17 The first example network structure of an NN filter is shown; Figure 18 A second example network structure for an NN filter is shown; Figure 19 The third example network structure of an NN filter is shown; Figure 20 The fourth example network structure of the NN filter is shown; Figure 21The fifth example network structure of an NN filter is shown; Figure 22 The sixth example network structure of an NN filter is shown; Figure 23 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 24 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0012] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation
[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0014] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0018] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.
[0023] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2The examples are shown separately.
[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.
[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0035] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0036] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0044] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0050] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0056] 1. Preliminary Discussion This disclosure relates to video encoding and decoding techniques. Specifically, it relates to loop filters in image / video encoding and decoding. It can be applied to existing video encoding and decoding standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), or standards yet to be finalized (e.g., AVS3). It can also be applied to future video encoding and decoding standards or video codecs, or used as a post-processing method outside the encoding / decoding process.
[0057] 2. Background Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, aiming to reduce the bitrate by 50% compared to HEVC. VVC version 1 was completed in July 2020.
[0058] 2.1. Color Space and Chromaticity Downsampling A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as tuples of numbers, typically 3 or 4 values or color components (e.g., RGB). Essentially, a color space is a refinement of a coordinate system and its subspaces.
[0059] For video compression, the most commonly used color spaces are YCbCr and RGB.
[0060] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with an apostrophe) is distinguished from Y, which is luminance, meaning that light intensity is encoded non-linearly based on gamma-corrected RGB primary colors.
[0061] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information. It takes advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences. 2.1.1. 4:4:4 Each of the three Y'CbCr components has the same sample rate, therefore there is no chromaticity downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2 The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by one-third, with little or no visual difference. 2.1.3. 4:2:0 In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating row. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variations of the 4:2:0 scheme with different horizontal and vertical positioning.
[0065] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (at the gap position).
[0066] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located at intervening positions, in the middle of alternating luminance samples.
[0067] In a 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternating rows.
[0068] 2.2. Definition of Video Unit The image is divided into one or more slice rows and one or more slice columns. A slice is a CTU sequence that covers a rectangular area of the image.
[0069] The sheet is divided into one or more bricks, each brick consisting of multiple CTU rows within the sheet.
[0070] A slice that is not divided into multiple bricks is also called a brick. However, a brick that is a proper subset of a slice is not called a slice.
[0071] A strip contains multiple slices of an image or multiple bricks of a slice.
[0072] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains multiple tiles that together form a rectangular area of the image. The tiles within the rectangular stripe are arranged in the order of the stripe's raster scan.
[0073] Figure 4 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips. Figure 4 In the image, which has 18×12 luminance CTUs, it is divided into 12 slices and 3 raster scan strips (informative).
[0074] In the VVC specification Figure 5 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips. Figure 5 In the image, which has 18×12 luminance CTUs, it is divided into 24 patches and 9 rectangular strips (informative).
[0075] In the VVC specification Figure 6 An example of an image divided into slices, bricks, and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the top left slice contains 1 brick, the top right slice contains 5 bricks, the bottom left slice contains 2 bricks, and the bottom right slice contains 3 bricks), and 4 rectangular strips. Figure 6 In the image, the image is divided into 4 pieces, 11 bricks, and 4 rectangular strips (for informational purposes).
[0076] 2.2.1. CTU / CTB Dimensions In VVC, the CTU size transmitted via signaling in SPS by the syntax element log2_ctu_size_minus2 can be as small as 4×4.
[0077] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[0078] The increment of 2 in log2_ctu_size_minus2 specifies the size of the luminance codec tree block for each CTU.
[0079] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma encoding / decoding block size.
[0080] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows: CtbLog2SizeY = log2_ctu_size_minus2+2(7-9) CtbSizeY = 1<<CtbLog2SizeY(7-10) MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2+2(7-11) MinCbSizeY = 1<<MinCbLog2SizeY(7-12) MinTbLog2SizeY = 2(7-13) MaxTbLog2SizeY = 6(7-14) MinTbSizeY = 1<<MinTbLog2SizeY(7-15) MaxTbSizeY = 1<<MaxTbLog2SizeY(7-16) PicWidthInCtbsY = Ceil(pic_width_in_luma_samples÷CtbSizeY)(7-17) PicHeightInCtbsY = Ceil(pic_height_in_luma_samples÷CtbSizeY)(7-18) PicSizeInCtbsY = PicWidthInCtbsY*PicHeightInCtbsY(7-19) PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY(7-20) PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY(7-21) PicSizeInMinCbsY = PicWidthInMinCbsY*PicHeightInMinCbsY(7-22) PicSizeInSamplesY = pic_width_in_luma_samples*pic_height_in_luma_samples(7-23) PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC(7-24) PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC(7-25) 2.2.2. CTUs in the Picture Assume that the CTB / LCU size is indicated by M×N (usually M equals N, as defined in HEVC / VVC), and for a CTB located at the boundary of a picture (or slice or strip or other kind of type, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For those CTBs depicted as in Figures 7A to 7C the CTB size still equals M×N, however, the lower boundary / right boundary of the CTB is outside the picture.
[0081] Figures 7A to 7C Examples of CTBs straddling the picture boundary are shown, Figure 7A showing K = M, L < N; Figure 7B showing K < M, L = N; Figure 7C showing K < M, L < N.
[0082] 2.3. Encoding and Decoding Processes of Typical Video Codecs Figure 8 An example of the encoder block diagram of VVC is shown, which includes three loop filter blocks: the deblocking filter (DF), sample adaptive offset (SAO), and ALF. Different from the DF that uses a predefined filter, SAO and ALF utilize the original samples of the current picture, by adding compensation and applying a finite impulse response (FIR) filter respectively, and reduce the mean square error between the original samples and the reconstructed samples by signaling compensation and filter coefficients using the encoded side information. ALF is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts caused by the previous stages.
[0083] 2.4. Deblocking Filter (DB) The input to the DB is the reconstructed sample before the loop filter.
[0084] The vertical edges in the image are first filtered. Then, the horizontal edges in the image are filtered using samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTB of each CTU are processed separately on a codec unit basis. The vertical edges of the codec blocks in the codec unit are filtered, starting from the left-hand edge of the codec block and proceeding geometrically through the edges towards the right-hand side of the codec block. The horizontal edges of the codec blocks in the codec unit are filtered, starting from the top edge of the codec block and proceeding geometrically through the edges towards the bottom of the codec block.
[0085] Figure 9 The image samples and horizontal and vertical block boundaries on an 8×8 grid are shown, as well as non-overlapping blocks of 8×8 samples that can be de-blocked in parallel.
[0086] 2.4.1. Boundary Determination The filter is applied to the 8×8 block boundaries. Additionally, the boundaries must be transform block boundaries or codec sub-block boundaries (e.g., due to affine motion prediction, ATMVP usage). For boundaries that are not such, the filter is disabled.
[0087] 2.4.2. Boundary Strength Calculation For transform block boundaries / encoder / decoder sub-block boundaries, if it lies within an 8x8 grid, it can be filtered, and the bS[xD] of that edge... i ][yD j ] (where [xD i ][yD j The settings for (representing coordinates) are defined in Table 1 and Table 2, respectively.
[0088] Table 1. Boundary Strength (when SPS IBC is disabled)
[0089] Table 2. Boundary Strength (when SPS IBC is enabled)
[0090] 2.4.3. Deblocking Decision for Luminance Component This section describes the process of deciding which square to remove. Figure 10 The pixels involved in the filter on / off decision and strong / weak filter selection are shown.
[0091] A wider and stronger brightness filter is used only when conditions 1, 2 and 3 are all true.
[0092] Condition 1 is the "large block condition". This condition detects whether the samples on the P-side and Q-side belong to a large block, and is represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0093] bSidePisLargeBlk = ((Edge type is vertical and p0 belongs to CU with width >= 32) || (Edge type is horizontal and p0 belongs to CU with height >= 32)) ? True : False bSideQisLargeBlk = ((Edge type is vertical and q0 belongs to CU with width >= 32) || (Edge type is horizontal and q0 belongs to CU with height >= 32)) ? True : False Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows: condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk) ? True : False Next, if condition 1 is true, condition 2 will be further examined. First, the following variables are derived: Following the method in HEVC, dp0, dp3, dq0, and dq3 are first derived. –if (p side is greater than or equal to 32) dp0 = ( dp0 + Abs( p50- 2 * p40+ p30) + 1 )>>1 dp3 = ( dp3 + Abs( p53- 2 * p43+ p33) + 1 )>>1 –if (q side is greater than or equal to 32) dq0 = ( dq0 + Abs( q50- 2 * q40+ q30) + 1 )>>1 dq3 = ( dq3 + Abs( q53- 2 * q43+ q33) + 1 )>>1 Condition 2 = (d < β) ? True : False Where d = dp0 + dq0 + dp3 + dq3. If conditions 1 and 2 are valid, then further check whether any block in the block uses a sub-block: If (bSidePisLargeBlk) { If (block P's mode == SUBBLOCKMODE) Sp = 5 else Sp = 7 } else Sp = 3 If (bSideQisLargeBlk) { If (block Q's mode == SUBBLOCKMODE) Sq = 5 else Sq = 7 } else Sq = 3 Finally, if both conditions 1 and 2 are valid, the proposed deblocking method will check condition 3 (the large block strong filter condition), which is defined as follows.
[0094] In condition 3 StrongFilterCondition, the following variables are deduced: Derive dpq in the manner described in HEVC.
[0095] Derive sp3 = Abs(p3 - p0) using the method described in HEVC. if (p side is greater than or equal to 32) if (Sp == 5) sp3= ( sp3+ Abs( p5- p3) + 1)>>1 else sp3= ( sp3+ Abs( p7- p3) + 1)>>1 Following the method in HEVC, we can derive sq3 = Abs(q0 - q3). if (q side is greater than or equal to 32) If (Sq == 5) sq3= ( sq3+ Abs( q5- q3) + 1)>>1 else sq3= ( sq3+ Abs( q7- q3) + 1)>>1 Following the HEVC convention, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (3 * β >> 5), and Abs(p0 - q0) < (5 * t) C + 1 )>>1) ? True: False.
[0096] 2.4.4. A more robust deblocking filter for luminance (designed for larger blocks) A bilinear filter is used when samples on either side of the boundary belong to a large block. Samples belonging to a large block are defined as those with a vertical edge width >= 32 and a horizontal edge height >= 32.
[0097] Bilinear filters are listed below.
[0098] Then HEVC is used to extract the block boundary sample points p in the block (as described above). i (i=0 to Sp-1) and q i (j=0 to Sq-1) (pi and qi are the i-th sample in the row used to filter vertical edges, or the i-th sample in the column used to filter horizontal edges) are replaced by the following linear interpolation: —
[0099] —
[0100] in The term is the position-dependent limiting described in Section 1.4.7, and The following is given.
[0101] 2.4.5. Deblocking control for chroma A strong chromaticity filter is used on both sides of the block boundary. Here, the chromaticity filter is selected when both sides of the chromaticity edge are greater than or equal to 8 (chromaticity position), and the following decisions with three conditions are satisfied: the first is for boundary strength and the decision of the block size. The proposed filter can be applied when the block width or height orthogonally spanning the block edge in the chromaticity sample domain is equal to or greater than 8. The second and third are essentially the same as those used for HEVC luminance deblocking decisions, which are the on / off decision and the strong filter decision, respectively.
[0102] In the first decision, the boundary strength (bS) is modified for chroma filtering, and the conditions are checked sequentially. If a condition is met, the remaining conditions with lower priority are skipped.
[0103] Chromatic deblocking is performed when bS equals 2, or when bS equals 1 when a large block boundary is detected.
[0104] The second and third conditions are essentially the same as those for the HEVC luminance filter determination below.
[0105] In the second condition: Then, based on the HEVC brightness, the cube is deduced, and d is derived.
[0106] The second condition will be true when d is less than β.
[0107] In the third condition, the strong filter condition is derived as follows: Derive dpq in the manner described in HEVC.
[0108] Derive sp3 = Abs(p3 - p0) using the method described in HEVC. Following the method in HEVC, we can derive sq3 = Abs(q0 - q3). According to the HEVC design, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (β >> 3), and Abs(p0 - q0) < (5 * t) C + 1 )>>1).
[0109] 2.4.6. Strong Deblocking Filter for Chroma Define the following strong deblocking filter for chroma: p2′= (3*p3+2*p2+p1+p0+q0+4)>>3 p1′= (2*p3+p2+2*p1+p0+q0+q1+4)>>3 p0′= (p3+p2+p1+2*p0+q0+q1+q2+4)>>3. The proposed chromaticity filter performs deblocking on a 4×4 chromaticity sample grid.
[0110] 2.4.7. Location-Related Limiting Position-dependent limiting (tcPD) is applied to the output samples of a brightness filtering process involving modifications to 7, 5, and 3 samples at the boundaries of a strong, long filter. Assuming a quantization error distribution, it is proposed to increase the limiting value for samples expected to have higher quantization noise, thus anticipating a higher deviation between the reconstructed sample values and the true sample values.
[0111] For each P or Q boundary filtered using an asymmetric filter, the position-related threshold table is selected from the two tables (i.e., Tc7 and Tc3 listed below) provided to the decoder as side information, based on the decision-making process in Section 1.4.2: Tc7 = { 6, 5, 4, 3, 2, 1, 1};Tc3 = { 6, 4, 2}; tcPD = (Sp == 3) ? Tc3 : Tc7; tcQD = (Sq == 3) ? Tc3 : Tc7; For P-boundaries or Q-boundaries filtered using short symmetric filters, a lower-amplitude position-related threshold is applied: Tc3 = { 3, 2, 1}; After defining the threshold, the filtered p’ i and q’ i The sample values are limited based on the tcP and tcQ limiting values: p’’ i = Clip3(p' i + tcP i , p’ i – tcP i , p’ i ); q’’ j = Clip3(q' j + tcQ j , q’ j – tcQ j , q’ j ); in p’ i and q’ i These are the filtered sample values. p’’ i and q’’ j It is the output sample value after amplitude limiting, and tcP i tcQ i From VVC tc parameters and tcPD and tcQD The derived clipping threshold. The function Clip3 is the clipping function as specified in VVC.
[0112] 2.4.8. Sub-block Removal and Adjustment To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying a maximum of 5 samples on the side using sub-block deblocking (AFFINE, ATMVP, or DMVR), as shown in the brightness control for long filters. Additionally, sub-block deblocking is adjusted such that sub-block boundaries on the 8×8 grid near the CU or implicit TU boundaries are restricted to modifying a maximum of two samples on each side.
[0113] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0114] If (block Q's mode == SUBBLOCKMODE && edge != 0) { if (!(implicitTU&&(edge == (64 / 4)))) if (edge == 2 || edge == (orthogonalLength - 2) || edge == (56 / 4) || edge == (72 / 4)) Sp = Sq = 2; else Sp = Sq = 3; else Sp = Sq = bSideQisLargeBlk ? 5:3 } Where edge = 0 corresponds to the CU boundary, edge = 2 or orthogonalLength-2 corresponds to the sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TU is used, then implicit TU is true.
[0115] 2.5.SAO The input to SAO is the reconstructed samples after DB. The concept of SAO is to reduce the average sample distortion of a region by first classifying region samples into multiple categories using a selected classifier, obtaining compensation for each category, and then adding the compensation to each sample in each category. The region's classifier index and compensation are encoded and decoded in the bitstream. In HEVC and VVC, a region (the unit used for SAO parameter signaling) is defined as a CTU.
[0116] HEVC employs two SAO types that meet the requirements of low complexity. These two types are Edge Compensation (EO) and Band Compensation (BO), which will be discussed in more detail below. The index of the SAO type is encoded and decoded (it is in the range of [0, 2]). For EO, sample classification is based on the comparison between the current sample and its neighboring samples according to the 1-D direction pattern: horizontal, vertical, 135° diagonal, and 45° diagonal.
[0117] Figures 11A to 11D Example diagrams are shown illustrating four 1-D orientation patterns used for EO sample point classification, namely horizontal (EO class = 0), vertical (EO class = 1), 135° diagonal (EO class = 2), and 45° diagonal (EO class = 3).
[0118] For a given EO class, each sample point within the CTB is classified into one of five categories. The current sample point value (labeled "c") is compared with its two neighboring samples along the selected 1-D pattern. The classification rules for each sample point are summarized in Table I. Categories 1 and 4 are associated with local valleys and local peaks along the selected 1-D pattern, respectively. Categories 2 and 3 are associated with concave and convex angles along the selected 1-D pattern, respectively. If the current sample point does not belong to EO categories 1 through 4, it is classified as category 0, and SAO is not applied.
[0119] Table 3: Sampling classification rules for edge compensation
[0120] 2.6. Adaptive Loop Filter Based on Geometric Transformation in JEM The input to DB is the reconstructed samples after DB and SAO. The sample classification and filtering processes are based on the reconstructed samples after DB and SAO.
[0121] In JEM, a geometrically transform-based adaptive loop filter (GALF) with block-based filter adaptation is applied. For the luminance component, one of 25 filters is selected for each 2×2 block based on the direction and activity of the local gradient.
[0122] 2.6.1. Filter Shape In JEM, there are a maximum of three diamond filter shapes (e.g. Figures 12A to 12C (As shown) can be selected for the luminance component. Indices are transmitted via signaling at the image level to indicate the filter shape used for the luminance component. Each square represents a sample point, and Ci (i = 0~6 (left), 0~12 (middle), 0~20 (right)) represents the coefficient to be applied to the sample point. For the chrominance component in the image, a 5×5 rhombus shape is always used.
[0123] Figures 12A to 12C An example of the shape of a GALF filter is shown ( Figure 12A : 5×5 rhombus Figure 12B : 7×7 rhombus Figure 12C (9×9 rhombus)
[0124] 2.6.1.1. Block Classification Each The block is categorized into one of 25 classes. (Category Index) C Based on its directionality and activity quantification value The derivation is as follows:
[0125] In order to calculate First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using 1-D Laplacian:
[0126] index Reference The coordinates of the sample point in the upper left corner of the block, and Indicator coordinates Reconstructed sample points at the location.
[0127] Then the gradients in the horizontal and vertical directions The maximum and minimum values are set as follows:
[0128] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0129] In order to derive directionality The values, these values are compared with each other and with two thresholds. Comparison: Step 1: If If both are true, then Set as .
[0130] Step 2: If If yes, continue from step 3; otherwise, continue from step 4.
[0131] Step 3: If ,but Set as ;otherwise Set as .
[0132] Step 4: If ,but Set as ;otherwise Set as .
[0133] Activity value Calculated as:
[0134] It is further quantized to the range of 0 to 4 (including boundary values), and the quantized value is represented as .
[0135] For the two chromaticity components in the image, no classification method is applied; that is, a single set of ALF coefficients is applied to each chromaticity component.
[0136] 2.6.1.2. Geometric Transformation of Filter Coefficients Figures 13A to 13C An example diagram is shown illustrating an example of a relative coordinator supported by a 5×5 rhombus filter.
[0137] Before filtering each 2×2 block, geometric transformations (such as rotation or diagonal and vertical flips) are applied to the coordinates, depending on the gradient values calculated for that block. k , l Associated filter coefficients This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make the different blocks more similar by aligning the directions of the different blocks to which the ALF is applied.
[0138] Three geometric transformations are introduced: diagonal, vertical flip, and rotation.
[0139] in It is the size of the filter, and These are coefficient coordinates, which make the position... In the top left corner, and in position In the bottom right corner. Depending on the gradient values computed for the block, the transformation is applied to the filter coefficients. f ( k , l The relationship between the transformation and the four gradients in the four directions is summarized in Table 4. Figure 12 shows the transformation coefficients at each position based on the 5×5 rhombus.
[0140] Table 4: Mapping of gradients to transformations computed for a block
[0141] 2.6.1.3. Filter Parameter Signaling In JEM, GALF filter parameters are transmitted via signaling for the first CTU, i.e., after the stripe header and before the SAO parameters of the first CTU. Up to 25 groups of luminance filter coefficients can be transmitted via signaling. To reduce bit overhead, filter coefficients from different categories can be merged. Furthermore, GALF coefficients from a reference image are stored and can be reused as GALF coefficients for the current image. The current image can optionally use GALF coefficients stored for a reference image, and GALF coefficient signaling is bypassed. In this case, only an index to one of the reference images is transmitted via signaling, and the stored GALF coefficients of the indicated reference image are inherited for the current image.
[0142] To support GALF temporal prediction, a candidate list of GALF filter sets is maintained. The candidate list is empty when decoding a new sequence. After decoding an image, the corresponding filter set can be added to the candidate list. Once the size of the candidate list reaches the maximum allowed value (i.e., 6 in the current JEM), new filter sets overwrite the oldest sets in decoding order; that is, a first-in, first-out (FIFO) rule is applied to update the candidate list. To avoid duplication, a set can only be added to the list if the corresponding image does not use GALF temporal prediction. To support temporal scalability, multiple candidate lists of filter sets exist, and each candidate list is associated with a temporal layer. More specifically, each array assigned by the temporal layer index (TempIdx) can constitute a filter set from previously decoded images with a TempIdx equal to or less than k. For example, the k-th array is assigned to be associated with a TempIdx equal to k, and it contains only filter sets from images with a TempIdx less than or equal to k. After a specific image is encoded or decoded, the set of filters associated with that image will be used to update those arrays associated with TempIdx that are equal to or higher than TempIdx.
[0143] Temporal prediction of GALF coefficients is used for inter-frame encoding / decoding to minimize signaling overhead. For intra-frames, temporal prediction is not available, and a set of 16 fixed filters is assigned to each class. To indicate the use of fixed filters, a flag for each class is transmitted via signaling, and the index of the selected fixed filter is also transmitted via signaling if necessary. Even when a fixed filter is selected for a given class, the coefficients of an adaptive filter can still be sent for that class. In this case, the coefficients of the filter to be applied to the reconstructed image are the sum of two sets of coefficients.
[0144] The filtering process for the luminance component can be controlled at the CU level. A flag is transmitted via signal to indicate whether GALF is applied to the luminance component of the CU. For the chrominance component, whether GALF is applied is indicated only at the image level.
[0145] 2.6.1.4. Filtering process On the decoder side, when GALF is enabled for a block, each sample within the block... The filtered sample values are obtained ,in L Indicates the filter length. Represents the filter coefficients, and This represents the decoded filter coefficients.
[0146] (10) Figure 14 This example illustrates relative coordinates used in a 5×5 diamond filter, assuming the current sample point's coordinates are (i, j) and (0, 0). Sample points at different coordinates filled with the same color are multiplied by the same filter coefficient.
[0147] 2.7. Adaptive Loop Filter Based on Geometric Transformation (GALF) in VVC 2.7.1. GALF in VTM-4 In VTM 4.0, the filtering process of the adaptive loop filter is performed as follows: (11) Among the sample points These are the input samples. These are the filtered output samples (i.e., the filtering result), and This represents the filter coefficients. In practice, VTM 4.0 uses integer arithmetic to achieve fixed-point precision calculations: (12) in L This represents the filter length, and where... These are the filter coefficients with fixed-point precision.
[0148] Compared to the current design of GALF in JEM, the current design of GALF in VVC has the following main changes: 1) Adaptive filter shapes have been removed. Only 7×7 filter shapes are allowed for the luminance component, and 5×5 filter shapes are allowed for the chrominance component.
[0149] 2) The signaling of ALF parameters has been removed from the strip / picture level to the CTU level.
[0150] 3) Class index calculation is performed at a 4×4 level instead of a 2×2 level. Furthermore, the Laplacian calculation method for downsampling used for ALF classification, as proposed in JVET-L0147, is utilized. More specifically, it is not necessary to calculate the horizontal / vertical / 45-degree diagonal / 135-degree gradient for each sample point within a block. Instead, 1:2 downsampling is used.
[0151] 2.8. Nonlinear ALF in Current VVC 2.8.1. Filtering Restatement Equation (11) can be restated as follows without affecting encoding / decoding efficiency: (13) in These are the same filter coefficients as in equation (11) [except for] It is equal to 1 in equation (13), and equal to 1 in equation (11). )].
[0152] Using the filter formula above (13), VVC introduces nonlinearity to achieve a higher amplitude near the sample value by using a simple limiting function. ) and the current sample value being filtered ( When the difference is too large, the influence of neighboring sample values is reduced, thus making ALF more efficient.
[0153] More specifically, the ALF filter is modified as follows: (14) in It is a limiting function, and It is the limiting parameter, which depends on Filter coefficients. The encoder performs optimization to find the optimal values. .
[0154] In the JVET-N0242 implementation, a limiting parameter is specified for each ALF filter. For each filter coefficient, a limiting value is transmitted via signal. This means that for each luminance filter, up to 12 limiting values can be transmitted via signal in the bitstream, and for a chrominance filter, up to 6 limiting values can be transmitted via signal in the bitstream.
[0155] To limit signaling costs and encoder complexity, only four fixed values are used, which are the same for inter-frame stripes and intra-frame stripes.
[0156] Because the variance of local differences in luminance is typically higher than that of local differences in chrominance, two different sets are applied for the luminance and chrominance filters. A maximum sample value is also introduced in each set (1024 bit depth for a 10-bit filter), allowing clipping to be disabled unnecessarily.
[0157] Table 5 provides the set of limiting values used in the JVET-N0242 test. Four values were selected by dividing the entire range of luminance sample values (encoded and decoded on 10 bits) and chrominance ranges from 4 to 1024 in the logarithmic domain by approximately equal division.
[0158] More precisely, the brightness table for the limit value has been obtained using the following formula: AlfClip L Where M=2 10 And N=4. (15) Similarly, the colorimetric table for the limiting values is obtained using the following formula: AlfClip C Where M=2 10 N=4, A=4. (16) Table 5: Authorized Limit Values
[0159] The selected limiting value is encoded or decoded in the "alf_data" syntax element using the Golomb coding scheme corresponding to the index of the limiting values in Table 5 above. This coding scheme is the same as that used for the filter index.
[0160] 2.9. Loop Filters Based on Convolutional Neural Networks for Video Encoding and Decoding 2.9.1. Convolutional Neural Networks In deep learning, convolutional neural networks (CNNs, or ConvNets) are a type of deep neural network most commonly used for analyzing visual images. They have been very successful in image and video recognition / processing, recommender systems, image classification, medical image analysis, and natural language processing.
[0161] CNNs are a regularized version of multilayer perceptrons. Multilayer perceptrons typically refer to fully connected networks, where each neuron in one layer is connected to all neurons in the next layer. This "fully connectedness" makes them prone to overfitting data. Typical regularization methods involve adding some form of magnitude measurement of the weights to the loss function. CNNs take a different approach to regularization: they utilize hierarchical patterns in the data and combine smaller, simpler patterns with more complex ones. Therefore, CNNs are at the lower extreme in terms of both connectivity and complexity.
[0162] Compared to other image classification / processing algorithms, CNNs use relatively little preprocessing. This means the network learns manually designed filters, unlike in traditional algorithms. This independence from prior knowledge and human intervention in feature design is a major advantage.
[0163] 2.9.2. Deep Learning for Image / Video Encoding and Decoding Deep learning-based image / video compression generally falls into two categories: end-to-end compression purely based on neural networks and traditional frameworks enhanced by neural networks. The first category typically employs an autoencoder-like structure, implemented through convolutional neural networks or recurrent neural networks. While relying solely on neural networks for image / video compression avoids any manual optimization or design, compression efficiency can be unsatisfactory. Therefore, works in the second category use neural networks as auxiliary tools and enhance traditional compression frameworks by replacing or strengthening certain modules. In this way, they can inherit the advantages of highly optimized traditional frameworks. For example, fully connected networks have been proposed for intra-frame prediction in HEVC. Besides intra-frame prediction, deep learning has also been utilized to enhance other modules. For instance, early work replaced the loop filters in HEVC with convolutional neural networks and achieved promising results. Early work applied neural networks to improve arithmetic encoding / decoding engines.
[0164] 2.9.3. Loop Filtering Based on Convolutional Neural Networks In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, thus introducing distortion into the reconstructed frame. To mitigate this distortion, convolutional neural networks can be trained to learn the mapping from distorted frames to the original frames. In practice, training must be performed before deploying CNN-based loop filtering.
[0165] 2.9.3.1. Training The goal of the training process is to find the optimal values of the parameters, including the weights and biases.
[0166] Codecs (e.g., HM, JEM, VTM, etc.) are used to compress the training dataset to generate distorted reconstructed frames.
[0167] The reconstructed frames are then fed into the CNN, and the cost is computed using the CNN's output and the ground truth frames (original frames). Common cost functions include SAD (Sum of Absolute Differences) and MSE (Mean Squared Error). Next, the gradient of the cost with respect to each parameter is derived using the backpropagation algorithm. The parameter values can be updated using these gradients. This process is repeated until the convergence criterion is met. After training is complete, the derived optimal parameters are saved for use in the inference phase.
[0168] 2.9.3.2. Convolution Process During convolution, the filter moves across the image from left to right and from top to bottom, changing a column of pixels on horizontal movement and a row of pixels on vertical movement. The amount of movement of the filter across the input image is called the stride, and it is almost always symmetrical in the height and width dimensions. The default stride or multiple strides in both dimensions are (1,1) for the height and width movement.
[0169] In most deep convolutional neural networks, residual blocks are used as basic modules and stacked multiple times to build the final network, where in one example, such as Figure 15B As shown, residual blocks are obtained by combining convolutional layers, ReLU / PReLU activation functions, and convolutional layers.
[0170] Figure 15A The architecture of the proposed CNN filter is shown, where M represents the number of feature maps and N represents the number of samples in one dimension. Figure 15B It shows Figure 15A The construction of ResBlock (residual block) in the code.
[0171] 2.9.3.3. Reasoning During the inference phase, distorted reconstructed frames are fed into the CNN and processed by the CNN model, whose parameters have been determined during the training phase. The input samples to the CNN can be reconstructed samples before or after DB, or before or after SAO, or before or after ALF.
[0172] 3. Problem Current neural network-based loop filtering has the following problems: a) The same convolution is used for each input of the NN-based loop filter. However, different inputs can have different importance. Therefore, it is reasonable to assign different convolutions with different kernel sizes and channel numbers to each input.
[0173] b) Convolutions with kernel size K are widely used in neural network-based loop filters. However, they can be decomposed into a combination of multiple convolutions with smaller kernel sizes to reduce complexity.
[0174] c) Side information generated during compression can be used as additional input to improve the performance of NN-based loop filters. For example, predicted image, strip type, boundary strength, basic QP, strip QP, and IPB information can be used as side information.
[0175] d) Multi-scaling structures used in neural networks can help improve the performance of NN-based loop filters.
[0176] e) NN-based networks are used to design loop filters. However, non-adjacent information is not considered. Transformer-based networks can capture global information.
[0177] f) NN-based filters have been proposed to enhance reconstruction. However, for some video content, traditional filters can outperform NN-based filters. Therefore, combining traditional filters with NN filters is reasonable.
[0178] g) The parameters of the NN-based filter are fixed during training. However, parameters such as QP, inference size, and block expansion size should be adapted to various video content during inference.
[0179] 4. Detailed Solution The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.
[0180] One or more neural network (NN) filter models are trained as part of a loop filtering technique or a filtering technique used in the post-processing stage to reduce distortion generated during compression. Samples with different characteristics are processed by different NN filter models. This disclosure details how to design a unified NN filter model by feeding at least one indicator that can be associated with a quality level (e.g., QP or constant rate factor (CRF) value or bit rate) / strip type / encoding / decoding mode / encoded information as input to the NN filter.
[0181] It should be noted that the concept of unifying the NN model by using feed indicators as input to the NN process can also be extended to other NN-based encoding and decoding tools, such as NN-based intra-frame prediction, NN-based cross-component prediction, NN-based inter-frame prediction, NN-based super-resolution, NN-based motion compensation, and NN-based transform design. In the following example, we use NN-based filtering techniques as an example.
[0182] It should also be noted that the concept of feeding encoded / decoded information as input to the neural network (NN) process can be extended to non-NN-based encoding / decoding tools, such as non-NN-based intra-frame prediction, non-NN-based cross-component prediction, non-NN-based inter-frame prediction, non-NN-based super-resolution, non-NN-based motion compensation, and non-NN-based transform design. For example, non-NN-based encoding / decoding tools can use encoded / decoded information to classify samples to be filtered into different categories.
[0183] In this disclosure, the NN filter can be any kind of NN filter, such as a convolutional neural network (CNN) filter, a fully connected neural network filter, a transformer-based filter, or a recurrent neural network-based filter.
[0184] In the following discussion, a video unit can be a sequence, image, strip, slice, brick, sub-image, CTU / CTB, CTU / CTB line, one or more CU / CB, one or more CTU / CTB, one or more VPDU (Virtual Pipeline Data Unit), or a sub-region within an image / strip / slice / brick. A parent video unit represents a unit larger than the video unit. Typically, a parent unit will contain multiple video units. For example, when the video unit is a CTU, the parent unit can be a strip, a CTU line, multiple CTUs, etc.
[0185] Simplifying NN filters through adaptive channel number 1. To solve problem 1, different convolution types are assigned to different inputs.
[0186] a. In one example, convolutions share the same kernel size, and different numbers of channels can be assigned for each input.
[0187] i. In one example, the number of channels for each input is represented as C1, C2, ..., C n , where n is an integer representing the number of inputs.
[0188] 1) In one example, additionally, a constraint is added such that for any index i and j (where i ∈ i ∈ j ... ), making C i C j .
[0189] 2) In one example, in addition, a constraint is added such that there exists at least one pair of indices i and j (where i and j are indices i and j). ), making C i C j .
[0190] b. In one example, convolutions share the same number of convolution channels, and different kernel sizes are assigned for each input.
[0191] i. In one example, Kernel size can be used for convolutions of a portion of the input, and The kernel size is used for convolution of the remaining input. K represents an integer value greater than 1.
[0192] c. In one example, different numbers of convolution channels and different kernel sizes are assigned to each input.
[0193] By decomposing and simplifying the NN filter 2. To solve problem 2, A convolution can be decomposed into a combination of multiple convolutions with smaller kernel sizes. K represents an integer value greater than 1. These represent the number of input channels and the number of output channels of the convolution, respectively.
[0194] a. In one example, when a convolution is decomposed, the number of input channels and the number of output channels of the convolution may remain unchanged.
[0195] i. In one example, Convolution is decomposed into Convolution and Combinations of convolutions.
[0196] ii. In one example, Convolution is decomposed into Convolution and A combination of convolutions, and any activation layer can be placed after each convolution.
[0197] iii. In one example, Convolution is decomposed into Convolution and Combinations of convolutions.
[0198] iv. In one example, Convolution is decomposed into Convolution and A combination of convolutions, and any activation layer can be placed after each convolution.
[0199] b. In one example, when a convolution is decomposed, the number of input channels and the number of output channels of the convolution can be changed.
[0200] i. In one example, Convolution is decomposed into Convolution and Combinations of convolutions, where It is different Positive integers.
[0201] ii. In one example, Convolution is decomposed into Convolution and A combination of convolutions, and any activation layer can be placed after each convolution, where It is different Positive integers.
[0202] iii. In one example, Convolution is decomposed into Convolution and Combinations of convolutions, where It is different Positive integers.
[0203] iv. In one example, Convolution is decomposed into Convolution and A combination of convolutions, and any activation layer can be placed after each convolution, where It is different Positive integers.
[0204] c. In one example, part The convolution is decomposed.
[0205] d. In one example, all Convolutions are all decomposed.
[0206] Designing the side information of NN filters 3. To address problem 3, it is specified which side information will be used as additional input to the NN-based loop filter and how they will be processed.
[0207] a. In one example, edge information can be used as additional input to a neural network-based loop filter.
[0208] i. In one example, a striped QP can be used as additional input to a loop filter based on an neural network.
[0209] 1) In one example, in addition, stripeQP is implemented via sliceQP. MAX_QP is normalized, where the value of MAX_QP can be 63.
[0210] 2) In one example, the strip QP is first sliced or unfolded into a two-dimensional array with the same size as the video unit to be filtered.
[0211] ii. In one example, the basic QP can be used as an additional input to an NN-based loop filter.
[0212] 1) In one example, in addition, the basic QP is derived from baseQP. MAX_QP is normalized, where the value of MAX_QP can be 63.
[0213] 2) In one example, the basic QP is first sliced or unfolded into a two-dimensional array with the same size as the video unit to be filtered.
[0214] iii. In one example, the predicted image can be used as additional input to a loop filter based on a neural network.
[0215] iv. In one example, the strip type can be used as additional input to a loop filter based on an NN.
[0216] 1) In one example, the strip type can also be a binary value that indicates whether the image to be filtered is an intra-frame stripe.
[0217] 2) In one example, the strip type indicator is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0218] v. In one example, the IPB information of the video unit to be filtered can be used as additional input to a NN-based loop filter.
[0219] 1) In one example, in addition, IPB information can be derived from the stripe type.
[0220] a) In one example, the value of the IPB information is derived as follows: if the stripe type is I stripe, the value of the IPB information is equal to a; if the stripe type is B stripe, the value of the IPB information is equal to b; if the stripe type is P stripe, the value of the IPB information is equal to c, where a, b, and c are constants.
[0221] i. In one example, b = c.
[0222] ii. In one example, b = -a, c = -a.
[0223] iii. In one example, a = 1, b = -1, c = -1.
[0224] iv. In one example, a = 1, b = 0.5, c = 0.5.
[0225] 2) In one example, the IPB information is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0226] vi. In one example, the boundary strength of the video cell to be filtered can be used as additional input to a neural network-based loop filter.
[0227] vii. In one example, any combination of the above side information can be used as additional input to an NN-based loop filter.
[0228] b. In one example, convolutions for each input edge information are performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image.
[0229] c. In one example, the reconstructed image and edge information are concatenated and then convolved.
[0230] 4. To address problem 4, multi-scale neural network structures can be used in NN-based loop filters.
[0231] a. In one example, a neural network with two branches can be used, and as follows: Figure 16 This illustrates one possible solution. Figure 16 In the middle, K i Indicates the core size (where ), C j Related to the number of channels (where ).
[0232] Design a converter-based NN filter 5. To address problem 5, a converter-based structure is proposed that can be incorporated into the design of NN filters.
[0233] a. In one example, the head / backbone / tail can be designed using a transformer network.
[0234] 6. To address problem 5, a transformer-based structure for designing NN filters can be combined with CNN.
[0235] a. In one example, in a network with NN filters, a CNN module can be followed by a transformer module.
[0236] 7. To address problem 5, it is proposed that transformer-based and CNN-based structures are alternatives in the design of NN filters.
[0237] a. In one example, when there are two NN-based filters, one containing a CNN-based filter and the other a transformer-based filter, the NN filter can be selected.
[0238] NN filter combined with DBF / SAO 8. To address problem 6, it was proposed that NN filters can be combined / fused / hybrid with traditional filters.
[0239] a. In one example, a traditional filter could be a deblocking filter (DBF).
[0240] b. In one example, a conventional filter could be Sample Adaptive Compensation (SAO).
[0241] c. In one example, a traditional filter could be a combination of DBF and SAO.
[0242] d. In one example, NN filters and DBFs can be fused / mixed.
[0243] i. In one example, SAO can be performed after the fusion / mixing of NN filters and DBF filters.
[0244] 1) In one example, SAO can be enabled / disabled by syntax elements in SPS / PPS, etc.
[0245] ii. In one example, ALF can be used after SAO.
[0246] 1) In one example, ALF can be enabled / disabled by syntax elements in SPS / PPS, etc.
[0247] e. In one example, NN filters, as well as combinations of DBF and SAO, can be fused / mixed.
[0248] i. In one example, an adaptive loop filter (ALF) can be used after fusing / mixing a combination of an NN filter and a DBF with a SAO.
[0249] 1) In one example, ALF can be enabled / disabled by syntax elements in SPS / PPS, etc.
[0250] Design scaling factor 9. To address problem 6, it is proposed that the reconstructed samples of traditional filters and NN filters be fused / mixed by a scaling factor.
[0251] a. In one example, reconstructed samples can be merged / blended at the strip / block level.
[0252] b. In one example, the scaling factor can be adaptive, and the scaling factor can be determined by the video content.
[0253] c. In one example, the scaling factor can be predefined.
[0254] d. In one example, the scaling factor can be separate for different components.
[0255] i. In one example, the scaling factor can be separate for the luminance and chrominance components.
[0256] ii. In one example, the scaling factor can be separate for the U component and the V component.
[0257] Design QP adjustment, number of parameters, and block QP adaptiveness. 10. To address problem 7, a candidate list containing multiple input parameters is proposed.
[0258] a. In one example, the number of candidates in the list can be configured at the sequence level / strip level / block level.
[0259] b. In one example, the candidate list can be constructed at the sequence level / strip level / block level.
[0260] c. In one example, the candidate list can be adaptive for sequence level / strip level / block level.
[0261] d. In one example, the input parameters depend on the representation as q The basic QP variables.
[0262] i. In one example, q It is a sequence-level / strip-level / block-level QP.
[0263] e. In one example, the candidate list may include a basic QP and an adjusted QP.
[0264] i. In one example, the adjusted QP can be achieved by adding an offset equal to the base QP. q .
[0265] ii. In one example, the offset can be equal to -5, 10, or 5.
[0266] iii. In one example, the offset may depend on the TID level.
[0267] Design adaptive inference granularity / configurable inference size 11. To address problem 7, it is proposed that the inference granularity / size of NN filters can be adaptive for sequence level / strip level / block level.
[0268] a. In one example, granularity / size can depend on the strip type.
[0269] b. In one example, the granularity / size can depend on the syntax elements at the sequence level / strip level / block level, which are configurable.
[0270] Design block expansion size 12. To address problem 7, it is proposed that block expansion / padding size can be adaptive at the sequence level / strip level / block level.
[0271] a. In one example, the block expansion / fill size may depend on the block pattern and / or stripe type.
[0272] b. In one example, the block expansion / padding size can depend on the configurable syntax elements at the sequence level / strip level / block level.
[0273] Design a unified NN filter 13. A unified NN filter is proposed to be designed by all or part of the items mentioned, which should not be interpreted in a narrow sense.
[0274] 5. Examples 5.1. Example 1 Figure 17 The first example network structure of the NN filter, which includes item 3 side information and item 4 multi-scale NN structure, is shown.
[0275] 5.2. Example 2 Figure 18 The second example network structure is shown, including 1. Adaptive channel number, 3. Side information, and 4. Multi-scaling NN structure for NN filters. Convolutions (Conv) A, B, ..., G are convolutions with different numbers of channels. Larger text fonts mean a larger number of channels; for example, convolution A is the convolution with the largest number of channels.
[0276] 5.3. Example 3 Figure 19 A third example network structure is shown, which includes item 2 decomposition, item 3 side information, and item 4 multi-scale NN structure for NN filters.
[0277] 5.4. Example 4 Figure 20 The fourth example network structure is shown, including NN filters with Item 1 adaptive channel number, Item 2 decomposition, Item 3 side information, and Item 4 multi-scaling NN structure. Convolutions A, B, ..., G are convolutions with different numbers of channels. Larger text fonts mean a larger number of channels; for example, convolution A is the convolution with the largest number of channels.
[0278] 5.5. Example 5 Figure 21 The fifth example network structure is shown, which includes 1. Adaptive channel number, 2. Decomposition, 3. Side information, 4. Multi-scaling NN structure, and 5. Transformer NN filter.
[0279] 5.6. Example 6 Figure 22 The sixth example network structure is shown, including Item 1 Adaptive Channel Number, Item 2 Decomposition, Item 3 Side Information, Item 4 Multi-Scaling NN Structure, and Item 5 Transformer NN Filter. Convolution A, Convolution B, ..., Convolution G are convolutions with different numbers of channels. Larger text fonts mean a larger number of channels; for example, Convolution A is the convolution with the largest number of channels.
[0280] As used herein, the term “video unit” or “video block” can be a sequence, picture, strip, slice, brick, subpicture, codec tree unit (CTU) / codec tree block (CTB), CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTU / CTB, one or more virtual pipeline data units (VPDU), or a sub-region within a picture / strip / slice / brick.
[0281] Figure 23 A flowchart of a method 2300 for video processing according to an embodiment of the present disclosure is shown. Method 2300 is implemented during the conversion between video units of a video and a bitstream of a video.
[0282] At box 2310, during the conversion between video units and video unit bitstreams, the neural network filter is determined according to rules. The rules indicate at least one of the following: different convolution types are assigned to different inputs to the neural network filter; convolutions with kernel sizes are decomposed into combinations of multiple convolutions with smaller kernel sizes; which side information is to be used as input to the neural network filter; multi-scaling neural network structures are used in the neural network filter; transformer-based structures are used in the neural network filter; non-neural network filters are combined with neural network filters; or a set of parameters for the neural network filter is adaptive.
[0283] At block 2320, a neural network filter is applied to the video unit. At block 2330, a conversion based on the filtered video unit is performed. In some embodiments, the conversion may include encoding the video unit into a bitstream. Alternatively or additionally, the conversion may include decoding the video unit from the bitstream. Compared to conventional solutions that directly select filters, filters can be adaptively combined for the video unit. In this way, encoding / decoding efficiency and effectiveness can be improved.
[0284] In some embodiments, convolutions sharing the same kernel size but with different numbers of channels are assigned to each input. In some embodiments, the number of channels for each input is represented as C1, C2, ..., C... n , where n is an integer representing the number of inputs.
[0285] In some embodiments, different constraints are added for different numbers of channels for different inputs. For example, the constraint for indices i and j is C. i C j , and among them Make C i C j .
[0286] In some embodiments, constraints are added that at least two inputs have different numbers of channels. For example, the constraint is that there exists at least one pair of indices i and j, C i C j , and among them .
[0287] In some embodiments, convolutions sharing the same number of channels but with different kernel sizes are assigned to each input. The kernel size is used for convolution of part of the input, and The kernel size is used for convolution of the remaining inputs, where K is an integer value greater than 1. In some other embodiments, different numbers of convolution channels and different kernel sizes are assigned to each input.
[0288] In some embodiments, Convolution is decomposed into a combination of multiple convolutions with smaller kernel sizes. In this case, K represents an integer value greater than 1. These represent the number of input channels and the number of output channels of the convolution, respectively.
[0289] In some embodiments, if the convolution is decomposed, the number of input channels and the number of output channels of the convolution remain unchanged. In some embodiments, Convolution is decomposed into Convolution and Combinations of convolutions. In some other embodiments, Convolution is decomposed into Convolution and The convolutions are combined, and activation layers are placed after each convolution. In some other embodiments, Convolution is decomposed into Convolution and Combinations of convolutions. In some other embodiments, Convolution is decomposed into Convolution and The convolutions are combined, and activation layers are placed after each convolution.
[0290] In some embodiments, if the convolution is decomposed, the number of input channels and the number of output channels of the convolution are changed. In some embodiments, Convolution is decomposed into Convolution and Combinations of convolutions, where It is different A positive integer. In some other embodiments, Convolution is decomposed into Convolution and The convolutions are combined, and activation layers are placed after each convolution, where It is different Positive integers. In some other embodiments, Convolution is decomposed into Convolution and Combinations of convolutions, where It is different A positive integer. In some other embodiments, Convolution is decomposed into Convolution and The convolutions are combined, and activation layers are placed after each convolution, where It is different Positive integers.
[0291] In some embodiments, portion The convolution is decomposed. In some other embodiments, all Convolutions are all decomposed.
[0292] In some embodiments, side information is used as additional input to the neural network filter. For example, the strip quantization parameter (QP) is used as additional input to the neural network filter. In some embodiments, the strip QP is expressed by the formula sliceQP. MAX_QP is normalized, where the value of MAX_QP is 63. In some other embodiments, the strip QP is first sliced or unfolded into a two-dimensional array with the same size as the video unit to be filtered.
[0293] In some embodiments, the base QP is used as additional input to the neural network filter. In some embodiments, the base QP is expressed by the formula baseQP. MAX_QP is normalized, where the value of MAX_QP is 63. In some other embodiments, the basic QP is first sliced or unfolded into a two-dimensional array with the same size as the video unit to be filtered.
[0294] In some embodiments, the predicted image is used as additional input to the neural network filter. In some other embodiments, the strip type is used as additional input to the neural network filter. For example, the strip type is a binary value that indicates whether the image to be filtered is an intra-frame stripe. As another example, the strip type indicator is first sliced or unfolded into a two-dimensional array with the same size as the video unit to be filtered.
[0295] In some embodiments, the IPB information of the video units to be filtered is used as additional input to the neural network filter. For example, the IPB information is derived from the strip type. In some embodiments, the values of the IPB information are derived as follows: if the strip type is I stripe, the value of the IPB information is equal to a; if the strip type is B stripe, the value of the IPB information is equal to b; and if the strip type is P stripe, the value of the IPB information is equal to c, where a, b, and c are constants. In some embodiments, b = c. Alternatively or additionally, b = -a, c = -a. Alternatively or additionally, a = 1, b = -1, c = -1. Alternatively or additionally, a = 1, b = 0.5, c = 0.5. In some embodiments, the IPB information is first sliced or unfolded into a two-dimensional array with the same size as the video units to be filtered.
[0296] In some embodiments, the boundary strength of the video unit to be filtered is used as additional input to the neural network filter. In some embodiments, a combination of side information is used as additional input to the neural network filter.
[0297] In some embodiments, convolutions for each input side information are performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image. In some embodiments, the reconstructed image and side information are concatenated, and then convolution is performed.
[0298] In some embodiments, the multi-scaling neural network architecture includes a used neural network with two branches. In some embodiments, such as Figure 16 As shown, one of the two branches includes Convolutional and activation layers, another branch includes Convolutional and activation layers, where K i Indicates the core size (where ), C j Related to the number of channels (where ).
[0299] In some embodiments, at least one of the head, backbone, or tail is determined using a transformer network. In some other embodiments, in a neural network filter, a transformer-based structure is combined with a convolutional neural network (CNN). For example, in a neural network filter, a CNN module is followed by a transformer module. In some further embodiments, both transformer-based and CNN-based structures are alternatives in a neural network filter. For example, if two neural network-based filters exist, including a CNN-based filter and a transformer-based filter, the neural network filter is selected.
[0300] In some embodiments, the non-neural network filter is a deblocking filter (DBF). In some other embodiments, the non-neural network filter is a sample adaptive compensation (SAO) filter. Alternatively, the non-neural network filter is a combination of a DBF and a SAO filter.
[0301] In some embodiments, the neural network filter and the DBF are combined. In some embodiments, the SAO filter follows the combination of the neural network filter and the DBF. For example, the SAO filter is enabled or disabled by a syntax element.
[0302] In some embodiments, an adaptive loop filter (ALF) is used after the SAO filter. In some embodiments, the ALF is enabled or disabled by a syntax element.
[0303] In some embodiments, a neural network filter and a DBF are combined with a SAO filter. For example, an ALF is used after combining a neural network filter and a DBF with a SAO filter. In some embodiments, the ALF is enabled or disabled by a syntax element.
[0304] In some embodiments, reconstructed samples from non-neural network filters and neural network filters are combined using a scaling factor. In some embodiments, the reconstructed samples are combined at the strip level or the block level. In some embodiments, the scaling factor is adaptive and determined by the video content. In some other embodiments, the scaling factor is predefined.
[0305] In some embodiments, the scaling factor is separate for different components. For example, the scaling factor is separate for the luminance component and the chrominance component. Alternatively or additionally, the scaling factor is separate for the U component and the V component.
[0306] In some embodiments, a candidate list including multiple input parameters is used. In some embodiments, the number of candidates in the candidate list is configurable at the sequence level, or the strip level, or the block level.
[0307] In some embodiments, the candidate list is constructed at the sequence level, or the strip level, or the block level. In some other embodiments, the candidate list is adaptive for the sequence level, or the strip level, or the block level.
[0308] In some embodiments, the input parameter is a variable that depends on a base QP denoted as q. For example, q is a QP at the sequence level, strip level, or block level.
[0309] In some embodiments, the candidate list includes a base QP and an adjusted QP. In some embodiments, the adjusted QP is equal to the base QP by adding an offset. In some embodiments, the offset is equal to one of the following: -5, 10, or 5. In some other embodiments, the offset depends on the thread identifier (TID) level.
[0310] In some embodiments, the inference granularity or size of the neural network filter is adaptive for one of the following: sequence level, strip level, or block level. In some embodiments, the inference granularity or size depends on the strip type. In some other embodiments, the inference granularity or size depends on whether a syntax element in one of the following is configurable: sequence level, strip level, or block level.
[0311] In some embodiments, at least one of the block expansion or padding size is adaptive to one of the following: sequence level, stripe level, or block level. In some embodiments, at least one of the block expansion or padding size depends on the block mode and / or stripe type. In some other embodiments, at least one of the block expansion or padding size depends on a syntax element that is configurable at the following: sequence level, stripe level, or block level.
[0312] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs of the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information is to be used as input to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with a neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; and generating a bitstream based on the filtered video units.
[0313] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs to the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information is to be used as input to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with a neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; generating a bitstream based on the filtered video units; and storing the bitstream in a non-transitory computer-readable recording medium.
[0314] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0315] Item 1. A video processing method comprising: during a conversion between video units of a video and a bitstream of the video units, determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs of the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information is to be used as inputs to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with the neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to the video units; and performing the conversion based on the filtered video units.
[0316] Item 2. The method according to Item 1, wherein convolutions sharing the same kernel size and different numbers of channels are assigned for each input.
[0317] Item 3. The method according to Item 2, wherein the number of channels for each input is represented as C1, C2, ..., C n , where n is an integer representing the number of inputs.
[0318] Item 4. The method according to Item 3, wherein a constraint is added that the number of channels is different for different inputs.
[0319] Item 5. According to the method described in Item 4, wherein the constraint, for indices i and j, is C i ≠C j And where 1≤ i , j≤ n .
[0320] Item 6. The method described in Item 3, wherein a constraint is added that at least two inputs have different numbers of channels.
[0321] Item 7. The method according to Item 6, wherein the constraint is that there exists at least one pair of indices i and j, C i ≠C j And where 1 ≤ i , j ≤ n .
[0322] Item 8. The method according to Item 1, wherein convolutions sharing the same number of channels and different kernel sizes are assigned for each input.
[0323] Item 9. The method according to Item 8, wherein a 1×1 kernel size is used for convolution of a portion of the input, and K × K The kernel size is used for convolution of the remaining input, where K is an integer value greater than 1.
[0324] Item 10. The method according to Item 1, wherein different numbers of convolution channels and different kernel sizes are assigned for each input.
[0325] Item 11. The method according to Item 1, wherein C 1× C 2× K × K The convolution is decomposed into the combination of the plurality of convolutions with smaller kernel sizes, where K represents an integer value greater than 1. C 1 and C 2 represents the number of input channels and the number of output channels of the convolution, respectively.
[0326] Item 12. The method according to Item 11, wherein if the convolution is decomposed, the number of input channels and the number of output channels of the convolution are not changed.
[0327] Item 13. The method according to Item 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2×1× K Convolution and C 1× C 2× K A combination of ×1 convolutions.
[0328] Item 14. The method according to Item 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2×1× K Convolution and C 1× C 2× K A combination of ×1 convolutions, with activation layers placed after each convolution.
[0329] Item 15. The method according to Item 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2× K ×1 convolution and C 1× C 2×1× K Combinations of convolutions.
[0330] Item 16. The method according to Item 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2× K ×1 convolution and C 1× C 2×1× K The convolutions are combined, and activation layers are placed after each convolution.
[0331] Item 17. The method according to Item 11, wherein if the convolution is decomposed, the number of input channels and the number of output channels of the convolution are changed.
[0332] Item 18. The method according to Item 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3×1× K Convolution and C 3× C 2× K A combination of ×1 convolutions, where C 3 is different C A positive integer of 2.
[0333] Item 19. The method according to Item 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3×1× K Convolution and C 3× C 2× K A combination of ×1 convolutions, with activation layers placed after each convolution, where C 3 is different C A positive integer of 2.
[0334] Item 20. The method according to Item 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3× K ×1 convolution and C 3× C 2×1× K Combinations of convolutions, where C 3 is different C A positive integer of 2.
[0335] Item 21. The method according to Item 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3× K ×1 convolution and C 3× C 2×1× K The convolutions are combined, and activation layers are placed after each convolution, where C 3 is different C A positive integer of 2.
[0336] Item 22. The method according to Item 11, wherein a portion thereof K × K The convolution is decomposed.
[0337] Item 23. The method according to Item 11, wherein all K × K Convolutions are all decomposed.
[0338] Item 24. The method according to Item 1, wherein the side information is used as additional input to the neural network filter.
[0339] Item 25. The method according to Item 24, wherein the strip quantization parameter (QP) is used as the additional input to the neural network filter.
[0340] Item 26. The method according to Item 25, wherein the stripe QP is obtained by the formula sliceQP. MAX_QP is normalized, with a value of 63.
[0341] Item 27. The method according to Item 25, wherein the strip QP is first sliced or unfolded into a two-dimensional array having the same size as the video unit to be filtered.
[0342] Item 28. The method according to Item 24, wherein the basic QP is used as an additional input to the neural network filter.
[0343] Item 29. The method according to Item 28, wherein the base QP is expressed by the formula baseQP. MAX_QP is normalized, with a value of 63.
[0344] Item 30. The method according to Item 28, wherein the basic QP is first sliced or unfolded into a two-dimensional array having the same size as the video unit to be filtered.
[0345] Item 31. The method according to Item 24, wherein the predicted image is used as additional input to the neural network filter.
[0346] Item 32. The method according to Item 24, wherein the strip type is used as an additional input to the neural network filter.
[0347] Item 33. The method according to Item 32, wherein the strip type is a binary value indicating whether the image to be filtered is an intra-frame stripe.
[0348] Item 34. The method according to Item 32, wherein the strip type indicator is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered.
[0349] Item 35. The method according to Item 24, wherein the IPB information of the video unit to be filtered is used as additional input to the neural network filter.
[0350] Item 36. The method according to Item 35, wherein the IPB information is derived from the stripe type.
[0351] Item 37. The method according to Item 36, wherein the value of the IPB information is derived as follows: if the stripe type is I stripe, then the value of the IPB information is equal to a; if the stripe type is B stripe, then the value of the IPB information is equal to b; and if the stripe type is P stripe, then the value of the IPB information is equal to c, where a, b, and c are constants.
[0352] Item 38. The method according to Item 37, wherein b = c, or wherein b = -a, c = -a, or wherein a = 1, b = -1, c = -1, or wherein a = 1, b = 0.5, c = 0.5.
[0353] Item 39. The method according to Item 35, wherein the IPB information is first sliced or unfolded into a two-dimensional array having the same size as the video unit to be filtered.
[0354] Item 40. The method according to Item 24, wherein the boundary intensity of the video unit to be filtered is used as additional input to the neural network filter.
[0355] Item 41. The method according to any one of items 24 to 40, wherein the combination of the side information is used as the additional input to the neural network filter.
[0356] Item 42. The method according to Item 1, wherein convolution for each input edge information is performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image.
[0357] Item 43. The method according to Item 1, wherein the reconstructed image and edge information are concatenated and then convolved.
[0358] Item 44. The method according to Item 1, wherein the multi-scaling neural network structure comprises a used neural network having two branches.
[0359] Item 45. The method according to Item 44, wherein one of the two branches includes C 1× C 2× K 1× K One branch consists of convolutional and activation layers, and another branch includes... C 3× C 4× K 2× K 2 convolutional and activation layers, where K i Denotes the kernel size, where 1 ≤ i ≤ 4, C j Related to the number of channels, where 1 ≤ j≤ 8.
[0360] Item 46. The method according to Item 1, wherein at least one of the head, backbone, or tail is determined by using a transformer network.
[0361] Item 47. The method according to Item 1, wherein in the neural network filter, the transformer-based structure is combined with a convolutional neural network (CNN).
[0362] Item 48. The method according to Item 47, wherein in the neural network filter, the CNN module is followed by a transformer module.
[0363] Item 49. The method according to Item 1, wherein the transformer-based structure and the CNN-based structure are alternatives in the neural network filter.
[0364] Item 50. The method according to Item 49, wherein if there are two neural network-based filters, including a CNN-based filter and a transformer-based filter, the neural network filter is selected.
[0365] Item 51. The method according to Item 1, wherein the non-neural network filter is a deblocking filter (DBF), or wherein the non-neural network filter is a sample adaptive compensation (SAO) filter, or wherein the non-neural network filter is a combination of a DBF and a SAO filter.
[0366] Item 52. The method according to Item 1, wherein the neural network filter and DBF are combined.
[0367] Item 53. The method according to Item 52, wherein the SAO filter follows the combination of the neural network filter and the DBF.
[0368] Item 54. The method according to Item 53, wherein the SAO filter is enabled or disabled by a syntax element.
[0369] Item 55. The method according to Item 52, wherein an adaptive loop filter (ALF) is used after the SAO filter.
[0370] Item 56. The method described in Item 55, wherein the ALF is enabled or disabled by a syntax element.
[0371] Item 57. The method according to Item 1, wherein the neural network filter and the combination of DBF and SAO filter are combined.
[0372] Item 58. The method according to Item 57, wherein the ALF is used after combining the neural network filter and the DBF with the SAO filter.
[0373] Item 59. The method described in Item 58, wherein the ALF is enabled or disabled by a syntax element.
[0374] Item 60. The method according to Item 1, wherein the reconstructed samples of the non-neural network filter and the reconstructed samples of the neural network filter are combined by a scaling factor.
[0375] Item 61. The method according to Item 60, wherein the reconstructed samples are combined at the strip level or the block level.
[0376] Item 62. The method according to Item 60, wherein the scaling factor is adaptive and determined by the video content.
[0377] Item 63. The method according to Item 60, wherein the scaling factor is predefined.
[0378] Item 64. The method according to Item 60, wherein the scaling factor is separate for different components.
[0379] Item 65. The method according to Item 64, wherein the scaling factor is separate for the luminance component and the chrominance component, and / or wherein the scaling factor is separate for the U component and the V component.
[0380] Item 66. The method described in Item 1 includes the use of a candidate list of multiple input parameters.
[0381] Item 67. The method according to Item 66, wherein the number of candidates in the candidate list is configurable at the sequence level, or the strip level, or the block level.
[0382] Item 68. The method according to Item 66, wherein the candidate list is constructed at the sequence level, or strip level, or block level.
[0383] Item 69. The method according to Item 66, wherein the candidate list is adaptive at the sequence level, or strip level, or block level.
[0384] Item 70. The method according to Item 66, wherein the input parameters are variables that depend on a basic QP denoted as q.
[0385] Item 71. The method according to Item 70, wherein q is a sequence-level, strip-level, or block-level QP.
[0386] Item 72. The method according to Item 66, wherein the candidate list includes a basic QP and an adjusted QP.
[0387] Item 73. The method according to Item 72, wherein the adjusted QP is equal to the base QP by adding an offset.
[0388] Item 74. The method according to Item 73, wherein the offset is equal to one of the following: -5, 10, 5.
[0389] Item 75. The method according to Item 73, wherein the offset depends on the thread identifier (TID) level.
[0390] Item 76. The method according to Item 1, wherein the inference granularity or size of the neural network filter is adaptive for one of the following: sequence level, strip level, or block level.
[0391] Item 77. The method according to Item 76, wherein the inference granularity or the size depends on the strip type.
[0392] Item 78. The method according to Item 76, wherein the inference granularity or the size depends on a syntax element that is configurable at one of the following levels: sequence level, strip level, or block level.
[0393] Item 79. The method according to Item 1, wherein at least one of the block expansion or padding size is adaptive for one of the following: sequence level, strip level, or block level.
[0394] Item 80. The method according to Item 79, wherein at least one of the block extension or padding size depends on the block pattern and / or stripe type.
[0395] Item 81. The method according to Item 79, wherein at least one of the block expansion or padding size depends on a syntax element that is configurable at the sequence level, strip level, or block level.
[0396] Item 82. The method according to any one of items 1 to 81, wherein the conversion includes encoding the video unit into the bitstream.
[0397] Item 83. The method according to any one of items 1 to 81, wherein the conversion includes decoding the video unit from the bitstream.
[0398] Item 84. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 83.
[0399] Item 85. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 83.
[0400] Item 86. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs of the neural network filter; a convolution having a kernel size is decomposed into a combination of multiple convolutions having a smaller kernel size; which side information is to be used as input to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with the neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; and generating the bitstream based on the filtered video units.
[0401] Item 87. A method for storing a bitstream of video, comprising: determining a neural network filter according to rules, wherein the rules indicate at least one of the following: different convolution types are assigned to different inputs of the neural network filter; a convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes; which side information is to be used as inputs to the neural network filter; a multi-scaling neural network structure is used in the neural network filter; a transformer-based structure is used in the neural network filter; a non-neural network filter is combined with the neural network filter; or a set of parameters of the neural network filter is adaptive; applying the neural network filter to video units of the video; generating the bitstream based on the filtered video units; and storing the bitstream in a non-transitory computer-readable recording medium.
[0402] Example device Figure 24 A block diagram of a computing device 2400 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2400 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0403] It should be understood that, Figure 24The computing device 2400 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0404] like Figure 24 As shown, computing device 2400 includes general-purpose computing device 2400. Computing device 2400 may include at least one or more processors or processing units 2410, memory 2420, storage unit 2430, one or more communication units 2440, one or more input devices 2450, and one or more output devices 2460.
[0405] In some embodiments, the computing device 2400 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2400 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0406] Processing unit 2410 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2420. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 2400. Processing unit 2410 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0407] Computing device 2400 typically includes various computer storage media. Such media can be any media accessible by computing device 2400, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2420 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2430 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2400.
[0408] The computing device 2400 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 24 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0409] Communication unit 2440 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2400 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 2400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0410] Input device 2450 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2460 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2440, computing device 2400 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2400 can also communicate with one or more devices that enable a user to interact with computing device 2400, or, if necessary, with any device that enables computing device 2400 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).
[0411] In some embodiments, some or all of the components of computing device 2400 may be arranged in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although to users they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0412] In embodiments of this disclosure, computing device 2400 may be used to implement video encoding / decoding. Memory 2420 may include one or more video encoding / decoding modules 2425 having one or more program instructions. These modules are accessible and executable by processing unit 2410 to perform the functions of the various embodiments described herein.
[0413] In an example embodiment of performing video encoding, input device 2450 may receive video data as input 2470 to be encoded. The video data may be processed, for example, by video codec module 2425 to generate an encoded bitstream. The encoded bitstream may be provided as output 2480 via output device 2460.
[0414] In an example embodiment of performing video decoding, input device 2450 may receive an encoded bitstream as input 2470. The encoded bitstream may be processed, for example, by video codec module 2425 to generate decoded video data. The decoded video data may be provided as output 2480 via output device 2460.
[0415] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A video processing method, comprising: During the conversion between video units and bitstreams of the video units, a neural network filter is determined according to rules, wherein the rules indicate at least one of the following: Different convolution types are assigned to different inputs to the neural network filters. A convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes. Which edge information will be used as input to the neural network filter? A multi-scaling neural network structure was used in the neural network filter. The converter-based structure is used in the neural network filter. The non-neural network filter is combined with the neural network filter, or The set of parameters of the neural network filter is adaptive; The neural network filter is applied to the video unit; as well as The conversion is performed based on the filtered video units.
2. The method of claim 1, wherein convolutions sharing the same kernel size and different numbers of channels are assigned for each input.
3. The method of claim 2, wherein the number of channels for each input is represented as C1, C2, ..., C... n , where n is an integer representing the number of inputs.
4. The method of claim 3, wherein a constraint is added that the number of channels is different for different inputs.
5. The method of claim 4, wherein the constraint, for indices i and j, is C. i ≠C j And where 1≤ i , j ≤ n .
6. The method of claim 3, wherein at least two inputs are constrained to have different numbers of channels.
7. The method of claim 6, wherein the constraint is the existence of at least one pair of indices i and j, C i ≠C j And where 1 ≤ i , j ≤ n .
8. The method of claim 1, wherein convolutions sharing the same number of channels and different kernel sizes are assigned for each input.
9. The method of claim 8, wherein a 1×1 kernel size is used for convolution of a portion of the input, and K × K The kernel size is used for convolution of the remaining input, where K is an integer value greater than 1.
10. The method of claim 1, wherein different numbers of convolution channels and different kernel sizes are assigned for each input.
11. The method of claim 1, wherein C 1× C 2× K × K The convolution is decomposed into the combination of the plurality of convolutions with smaller kernel sizes, where K represents an integer value greater than 1. C 1 and C 2 represents the number of input channels and the number of output channels of the convolution, respectively.
12. The method of claim 11, wherein if the convolution is decomposed, the number of input channels and the number of output channels of the convolution are not changed.
13. The method of claim 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2×1× K Convolution and C 1× C 2× K A combination of ×1 convolutions.
14. The method of claim 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2×1× K Convolution and C 1× C 2× K A combination of ×1 convolutions, with activation layers placed after each convolution.
15. The method of claim 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2× K ×1 convolution and C 1× C 2×1× K Combinations of convolutions.
16. The method of claim 11, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 2× K ×1 convolution and C 1× C 2×1× K The convolutions are combined, and activation layers are placed after each convolution.
17. The method of claim 11, wherein if the convolution is decomposed, the number of input channels and the number of output channels of the convolution are changed.
18. The method of claim 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3×1× K Convolution and C 3× C 2× K A combination of ×1 convolutions, where C 3 is different C A positive integer of 2.
19. The method of claim 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3×1× K Convolution and C 3× C 2× K A combination of ×1 convolutions, with activation layers placed after each convolution, where C 3 is different C A positive integer of 2.
20. The method of claim 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3× K ×1 convolution and C 3× C 2×1× K Combinations of convolutions, where C 3 is different C A positive integer of 2.
21. The method of claim 17, wherein... C 1× C 2× K × K Convolution is decomposed into C 1× C 3× K ×1 convolution and C 3× C 2×1× K The convolutions are combined, and activation layers are placed after each convolution, where C 3 is different C A positive integer of 2.
22. The method of claim 11, wherein a portion thereof K × K The convolution is decomposed.
23. The method of claim 11, wherein all K × K Convolutions are all decomposed.
24. The method of claim 1, wherein the side information is used as additional input to the neural network filter.
25. The method of claim 24, wherein the strip quantization parameter (QP) is used as the additional input to the neural network filter.
26. The method of claim 25, wherein the stripe QP is calculated using the formula sliceQP. MAX_QP is normalized, with a value of 63.
27. The method of claim 25, wherein the strip QP is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered.
28. The method of claim 24, wherein the basic QP is used as an additional input to the neural network filter.
29. The method of claim 28, wherein the base QP is defined by the formula baseQP. MAX_QP is normalized, with a value of 63.
30. The method of claim 28, wherein the basic QP is first sliced or unfolded into a two-dimensional array having the same size as the video unit to be filtered.
31. The method of claim 24, wherein the predicted image is used as additional input to the neural network filter.
32. The method of claim 24, wherein the strip type is used as additional input to the neural network filter.
33. The method of claim 32, wherein the stripe type is a binary value, the binary value indicating whether the image to be filtered is an intra-frame stripe.
34. The method of claim 32, wherein the strip type indicator is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered.
35. The method of claim 24, wherein the IPB information of the video unit to be filtered is used as additional input to the neural network filter.
36. The method of claim 35, wherein the IPB information is derived from the stripe type.
37. The method of claim 36, wherein the value of the IPB information is derived as follows: If the stripe type is I stripe, then the value of the IPB information is equal to a. If the stripe type is B stripe, then the value of the IPB information is equal to b, and If the stripe type is P-strip, then the value of the IPB information is equal to c. Where a, b, and c are constants.
38. The method of claim 37, wherein b = c, or Where b = -a, c = -a, or Where a = 1, b = -1, c = -1, or Where a = 1, b = 0.5, c = 0.
5.
39. The method of claim 35, wherein the IPB information is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered.
40. The method of claim 24, wherein the boundary strength of the video unit to be filtered is used as additional input to the neural network filter.
41. The method according to any one of claims 24 to 40, wherein the combination of side information is used as the additional input to the neural network filter.
42. The method of claim 1, wherein convolution for each input edge information is performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image.
43. The method of claim 1, wherein the reconstructed image and edge information are concatenated and then convolved.
44. The method of claim 1, wherein the multi-scale neural network structure comprises a used neural network having two branches.
45. The method of claim 44, wherein one of the two branches comprises C 1× C 2× K 1× K 1. Convolutional and activation layers, Another branch includes C 3× C 4× K 2× K 2 convolutional and activation layers, and Where K i Denotes the kernel size, where 1 ≤ i ≤ 4, C j Related to the number of channels, where 1 ≤ j ≤ 8.
46. The method of claim 1, wherein at least one of the head, trunk, or tail is determined by using a converter network.
47. The method of claim 1, wherein in the neural network filter, the transformer-based structure is combined with a convolutional neural network (CNN).
48. The method of claim 47, wherein in the neural network filter, the CNN module is followed by a transformer module.
49. The method of claim 1, wherein the transformer-based structure and the CNN-based structure are alternatives in the neural network filter.
50. The method of claim 49, wherein if there are two neural network-based filters, including a CNN-based filter and a transformer-based filter, the neural network filter is selected.
51. The method of claim 1, wherein the non-neural network filter is a deblocking filter (DBF), or The non-neural network filter mentioned above is a sample adaptive compensation (SAO) filter, or The non-neural network filter mentioned therein is a combination of DBF and SAO filters.
52. The method of claim 1, wherein the neural network filter and the DBF are combined.
53. The method of claim 52, wherein the SAO filter follows the combination of the neural network filter and the DBF.
54. The method of claim 53, wherein the SAO filter is enabled or disabled by a syntax element.
55. The method of claim 52, wherein an adaptive loop filter (ALF) is used after the SAO filter.
56. The method of claim 55, wherein the ALF is enabled or disabled by a syntax element.
57. The method of claim 1, wherein the neural network filter and the combination of DBF and SAO filter are combined.
58. The method of claim 57, wherein the ALF is used after the combination of the neural network filter and the DBF with the SAO filter is performed.
59. The method of claim 58, wherein the ALF is enabled or disabled by a syntax element.
60. The method of claim 1, wherein the reconstructed samples of the non-neural network filter and the reconstructed samples of the neural network filter are combined by a scaling factor.
61. The method of claim 60, wherein the reconstructed samples are combined at the strip level or the block level.
62. The method of claim 60, wherein the scaling factor is adaptive and determined by the video content.
63. The method of claim 60, wherein the scaling factor is predefined.
64. The method of claim 60, wherein the scaling factor is separate for different components.
65. The method of claim 64, wherein the scaling factor is separate for the luminance component and the chrominance component, and / or The scaling factor is separate for the U component and the V component.
66. The method of claim 1, wherein a candidate list of multiple input parameters is used.
67. The method of claim 66, wherein the number of candidates in the candidate list is configurable at the sequence level, or the strip level, or the block level.
68. The method of claim 66, wherein the candidate list is constructed at the sequence level, or the strip level, or the block level.
69. The method of claim 66, wherein the candidate list is adaptive at the sequence level, or strip level, or block level.
70. The method of claim 66, wherein the input parameters are variables that depend on a basic QP denoted as q.
71. The method of claim 70, wherein q is a sequence-level, strip-level, or block-level QP.
72. The method of claim 66, wherein the candidate list comprises a basic QP and an adjusted QP.
73. The method of claim 72, wherein the adjusted QP is equal to the base QP by adding an offset.
74. The method of claim 73, wherein the offset is equal to one of the following: -5, 10, 5.
75. The method of claim 73, wherein the offset depends on the thread identifier (TID) level.
76. The method of claim 1, wherein the inference granularity or size of the neural network filter is adaptive for one of the following: sequence level, strip level, or block level.
77. The method of claim 76, wherein the inference granularity or the size depends on the stripe type.
78. The method of claim 76, wherein the inference granularity or the size depends on a syntax element that is configurable at one of the following levels: sequence level, strip level, or block level.
79. The method of claim 1, wherein at least one of the block expansion or padding size is adaptive to one of the following: sequence level, strip level, or block level.
80. The method of claim 79, wherein at least one of the block extension or fill size depends on the block pattern and / or stripe type.
81. The method of claim 79, wherein at least one of the block expansion or padding size depends on a syntax element that is configurable at the sequence level, strip level, or block level.
82. The method according to any one of claims 1 to 81, wherein the conversion comprises encoding the video unit into the bitstream.
83. The method according to any one of claims 1 to 81, wherein the conversion comprises decoding the video unit from the bitstream.
84. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 83.
85. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 83.
86. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The neural network filter is determined according to rules, wherein the rules indicate at least one of the following: Different convolution types are assigned to different inputs to the neural network filters. A convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes. Which edge information will be used as input to the neural network filter? A multi-scaling neural network structure was used in the neural network filter. The converter-based structure is used in the neural network filter. The non-neural network filter is combined with the neural network filter, or The set of parameters of the neural network filter is adaptive; The neural network filter is applied to the video unit of the video; as well as The bitstream is generated based on the filtered video units.
87. A method for storing a bitstream of video, comprising: The neural network filter is determined according to rules, wherein the rules indicate at least one of the following: Different convolution types are assigned to different inputs to the neural network filters. A convolution with a kernel size is decomposed into a combination of multiple convolutions with smaller kernel sizes. Which edge information will be used as input to the neural network filter? A multi-scaling neural network structure was used in the neural network filter. The converter-based structure is used in the neural network filter. The non-neural network filter is combined with the neural network filter, or The set of parameters of the neural network filter is adaptive; The neural network filter is applied to the video unit of the video; The bitstream is generated based on the filtered video units; as well as The bitstream is stored in a non-transitory computer-readable recording medium.