Improved neural network-based super-resolution with respect to designing video coding and decoding
By introducing neural network-based super-resolution technology and filter optimization into video encoding and decoding, the problem of low coding efficiency in high-resolution video has been solved, achieving more efficient bandwidth utilization and improved image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video encoding and decoding technologies suffer from high bandwidth requirements and low coding efficiency when processing high-resolution videos, especially in color space downsampling and video unit division, where they struggle to effectively utilize the characteristics of the human visual system.
The super-resolution (NNSR) technology based on neural networks optimizes the encoding and decoding process of video data by determining the strip quantization parameter (QP) as an additional input. Combined with filter and classifier techniques, it improves coding efficiency and image quality.
It improves the encoding efficiency and image quality of video encoding and decoding, reduces bandwidth requirements, and enhances the flexibility and adaptability of video processing.
Smart Images

Figure CN121970080A_ABST
Abstract
Description
Neural network-based super-resolution for improving video encoding and decoding design
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to International Patent Application No. PCT / CN2023 / 123048, filed on September 29, 2023. The entire contents of the aforementioned patent application are incorporated herein by reference. Technical Field
[0003] This patent document relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining an application strip quantization parameter (QP) as an additional input to a neural network (NN)-based super-resolution (SR) process; and performing a conversion between visual media data and a bitstream based on the NN-based SR.
[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the preceding aspects.
[0007] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec device performs the method of any of the preceding aspects.
[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining an application of strip quantization parameters (QP) as additional input to a neural network (NN)-based super-resolution (SR) process; and generating the bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining an application strip quantization parameter (QP) as additional input to a neural network (NN)-based super-resolution (SR) process; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] For clarity, any of the embodiments described above may be combined with one or more other embodiments described above to create new embodiments within the scope of this disclosure.
[0011] These and other features will become clearer through the following detailed description in conjunction with the accompanying drawings and claims. Attached Figure Description
[0012] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.
[0013] Figure 1 shows an example image segmented into raster scan strips.
[0014] Figure 2 shows an example image segmented into rectangular scan strips.
[0015] Figure 3 shows an example image of the bricks.
[0016] Figures 4A-C show examples of codec tree blocks (CTBs) spanning image boundaries.
[0017] Figure 5 shows an example of an encoder block diagram.
[0018] Figure 6 shows an example of block boundaries in the image.
[0019] Figure 7 shows an example involving the pixels used by the filter.
[0020] Figure 8 shows an example of the orientation pattern for edge offset (EO) sample point classification.
[0021] Figure 9 shows an example of the shape of an adaptive loop filter (GALF) based on geometric transformation.
[0022] Figure 10 shows an example of relative coordinates supported by a 5×5 diamond filter.
[0023] Figure 11 shows an example of relative coordinates supported by a 5×5 diamond filter.
[0024] Figure 12A shows an example convolutional neural network (CNN) filter.
[0025] Figure 12B shows an example construction of the residual block.
[0026] Figure 13A shows an example network structure for neural network (NN) super-resolution (SR).
[0027] Figure 13B shows an example skeleton block.
[0028] Figure 14A shows an example network structure for NN SR.
[0029] Figure 14B shows an example skeleton block.
[0030] Figure 15 is a block diagram illustrating an example video processing system.
[0031] Figure 16 is a block diagram of an example video processing device.
[0032] Figure 17 is a flowchart of an example method for video processing.
[0033] Figure 18 is a block diagram illustrating an example video codec system.
[0034] Figure 19 is a block diagram showing an example encoder.
[0035] Figure 20 is a block diagram showing an example decoder.
[0036] Figure 21 is a schematic diagram of an example encoder. Detailed Implementation
[0037] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but modifications can be made within the scope of the appended claims, together with their entire equivalents.
[0038] The use of chapter headings in this document is for ease of understanding and not to limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter only. Furthermore, the technologies described herein are applicable to other video codec protocols and designs.
[0039] 1. Preliminary Discussion
[0040] This document relates to video codec technology. Specifically, it relates to super-resolution in image / video codecs. It can be applied to video codec standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), or the Audio-Video Standard 3 (AVS3). It can also be applied to other video codec technologies, video codecs, and / or used as a post-processing method outside of the encoding and decoding processes.
[0041] 2. Video codec standards
[0042] Video coding standards have evolved primarily through the development of standards by the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC). ITU-T produced the H.261 and H.263 standards, ISO / IEC produced the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly produced the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard [1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established by the Video Coding Experts Group (VCEG) and MPEG. JVET has adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM) [2]. The JVET between the Video Codec Experts Group (VCEG) (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (e.g., MPEG) was created in collaboration with the VVC standard, aiming to reduce the bit rate by 50% compared to HEVC.
[0043] A sample version of the VVC draft, namely the Multi-Functional Video Codec (Draft 10), can be found at: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=10399. A sample version of the VVC reference software, named VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags / VTM-10.0.
[0044] 2.1 Color Space and Chromaticity Downsampling
[0045] A color space, also known as a color model (or color system), is a mathematical model that describes a range of colors as tuples of numbers, such as 3 or 4 values or color components (e.g., RGB). Generally, a color space is a refinement of a coordinate system and its subspaces. For video compression, the most commonly used color spaces are Luminance, Blue Difference, and Red Difference (YCbCr) and Red, Green, and Blue (RGB).
[0046] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with an apostrophe) is distinguished from Y, which is luminance, meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0047] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences.
[0048] 2.1.1 4:4:4
[0049] In a 4:4:4 scheme, each of the three Y'CbCr components has the same sampling rate. Therefore, there is no chromaticity downsampling. This scheme is sometimes used in high-end cinema scanners and film post-production.
[0050] 2.1.2 4:2:2
[0051] In 4:2:2, the two chroma components are sampled at half the sampling rate of the luminance. The horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by one-third, resulting in almost no visual difference.
[0052] 2.1.3 4:2:0
[0053] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating row. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.
[0054] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (at gap positions). In Joint Picture Experts Group (JPEG) / JPEG File Exchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are located at gap positions, in the middle of alternating luminance samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternating lines.
[0055] 2.2 Definition of Video Unit
[0056] An image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. A slice can be divided into one or more bricks, each brick comprising some rows of code-decoder tree units (CTUs) within the slice. A slice that is not divided into multiple bricks can also be called a brick. However, bricks that are a proper subset of a slice cannot be called a slice. A strip contains multiple slices of an image or multiple bricks of a slice.
[0057] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains multiple bricks of the image that together form a rectangular area of the image. The bricks within the rectangular stripe are arranged in the order of the stripe's brick raster scan. Figure 1 shows an example of raster scan stripe segmentation for an image with 18×12 luminance CTUs, where the image is divided into 12 slices and 3 raster scan stripes.
[0058] Figure 2 shows an example of rectangular strip segmentation of an image with 18×12 luminance CTUs, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0059] Figure 3 shows an example of an image divided into slices, bricks, and rectangular strips, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the top left slice contains 1 brick, the top right slice contains 5 bricks, the bottom left slice contains 2 bricks, and the bottom right slice contains 3 bricks) and 4 rectangular strips.
[0060] 2.2.1 CTU / CTB Dimensions
[0061] In VVC, the CTU size transmitted via signaling in the Sequence Parameter Set (SPS) by the syntax element log2_ctu_size_minus2 can be as small as 4×4.
[0062] 7.3.2.3 Sequence Parameter Set Raw Byte Sequence Payload (RBSP) Syntax
[0063]
[0064]
[0065]
[0066] `log2_ctu_size_minus2` plus 2 specifies the luma codec block size for each CTU. `log2_min_luma_coding_block_size_minus2` plus 2 specifies the minimum luma codec block size. The variables `CtbLog2SizeY`, `CtbSizeY`, `MinCbLog2SizeY`, `MinCbSizeY`, `MinTbLog2SizeY`, `MaxTbLog2SizeY`, `MinTbSizeY`, `MaxTbSizeY`, `PicWidthInCtbsY`, `PicHeightInCtbsY`, `PicSizeInCtbsY`, `PicWidthInMinCbsY`, `PicHeightInMinCbsY`, `PicSizeInMinCbsY`, `PicSizeInSamplesY`, `PicWidthInSamplesC`, and `PicHeightInSamplesC` are derived as follows:
[0067] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0068] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0069] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0070] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0071] MinTbLog2SizeY = 2 (7-13)
[0072] MaxTbLog2SizeY = 6 (7-14)
[0073] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0074] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0075] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0076] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY ) (7-18)
[0077] PicSizeInCtbsY = PicWidthInCtbsY PicHeightInCtbsY (7-19)
[0078] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0079] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0080] PicSizeInMinCbsY = PicWidthInMinCbsY PicHeightInMinCbsY (7-22)
[0081] PicSizeInSamplesY = pic_width_in_luma_samples pic_height_in_luma_samples (7-23)
[0082] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0083] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0084] 2.2.2 CTU in the picture
[0085] Figure 4A shows an example of a CTB across the bottom image boundary. Figure 4B shows an example of a CTB across the right image boundary. Figure 4C shows an example of a CTB across the bottom right image boundary.
[0086] Assume that the CTB / maximum coding unit (LCU) size is indicated by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or stripe or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For those CTBs depicted in FIGS. 4A - 4C, the CTB size still equals M×N; however, the lower boundary / right boundary of the CTB is outside the picture.
[0087] 2.3 Coding and decoding processes of an example video codec
[0088] FIG. 5 shows an example of an encoder block diagram of VVC, which includes three loop filter blocks: a deblocking filter (DF), sample adaptive offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and use the coded side information to signal the offset and filter coefficients to reduce the mean square error between the original samples and the reconstructed samples. ALF is located in the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts caused by the previous stages.
[0089] 2.4 Deblocking filter (DB)
[0090] FIG. 6 shows an example of a block boundary in a picture. The input of the DB is the reconstructed samples before the loop filter. FIG. 6 shows an example of picture samples, horizontal block boundaries, and vertical block boundaries on an 8×8 grid, as well as non - overlapping blocks of 8×8 samples that can be deblocked in parallel.
[0091] The vertical edges in the picture are first filtered. Then, the horizontal edges in the picture are filtered with the samples modified by the vertical edge filtering process as the input. The vertical and horizontal edges in the CTB of each CTU are processed separately on the basis of the coding unit. The vertical edges of the coding blocks in the coding unit are filtered, starting from the edge on the left hand side of the coding block and proceeding through the edges in their geometric order towards the right hand side of the coding block. The horizontal edges of the coding blocks in the coding unit are filtered, starting from the edge on the top of the coding block and proceeding through the edges in their geometric order towards the bottom of the coding block.
[0092] 2.4.1 Boundary decision
[0093] Filtering is applied to the 8×8 block boundaries. Additionally, such boundaries must be transform block boundaries or coding sub - block boundaries, for example, due to the use of affine motion vector prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0094] 2.4.2 Boundary strength calculation
[0095] For transform block boundaries / encoder / decoder sub-block boundaries, if the boundary is located in an 8×8 grid, the boundary can be filtered, and the settings of bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) of the edge are defined as in Tables 1 and 2, respectively.
[0096] Table 1 Boundary Strength (when SPS IBC is disabled)
[0097]
[0098] Table 2 Boundary Strength (when SPS IBC is enabled)
[0099]
[0100] 2.4.3 Deblocking decision for the luminance component
[0101] This section describes the deblocking decision process. Figure 7 shows an example of pixels involved in filter usage. Figure 7 illustrates pixels involved in filter on / off decisions and strong / weak filter selection.
[0102] A wider and stronger brightness filter is used only when conditions 1, 2, and 3 are all true. Condition 1 is the "bulk condition." This condition detects whether samples on the P-side and Q-side belong to a bulk, and is represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0103] bSidePisLargeBlk = ((edge type is vertical and p0 belongs to a codec unit (CU) with width >= 32)) || (edge type is horizontal and p0 belongs to a CU with height >= 32)) ? TRUE: FALSE
[0104] bSideQisLargeBlk = ((Edge type is vertical and q0 belongs to CU with width >= 32) || (Edge type is horizontal and q0 belongs to CU with height >= 32)) ? TRUE: FALSE
[0105] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0106] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk) ? TRUE: FALSE
[0107] Next, if condition 1 is true, then condition 2 will be further checked. First, derive the following variables:
[0108] First, derive dp0, dp3, dq0, and dq3 using the HEVC method.
[0109] if (p side is greater than or equal to 32)
[0110] dp0 = (dp0 + Abs(p50 - 2)) p40 + p30 + 1) >> 1
[0111] dp3 = (dp3 + Abs(p53 - 2)) p43 + p33 + 1) >> 1
[0112] if (q side is greater than or equal to 32)
[0113] dq0 = (dq0 + Abs(q50 - 2)) q40 + q30 + 1) >> 1
[0114] dq3 = (dq3 + Abs(q53 - 2)) q43 + q33 + 1) >> 1
[0115] Condition 2 = (d < β) ? TRUE : FALSE
[0116] Where d = dp0 + dq0 + dp3 + dq3.
[0117] If conditions 1 and 2 are valid, then further check whether any of the blocks uses a sub-block:
[0118] If (bSidePisLargeBlk)
[0119] {
[0120] If (block P's mode == SUBBLOCKMODE)
[0121] Sp = 5
[0122] else
[0123] Sp = 7
[0124] }
[0125] else
[0126] Sp = 3
[0127] If (bSideQisLargeBlk)
[0128] {
[0129] If (block Q's mode == SUBBLOCKMODE)
[0130] Sq = 5
[0131] else
[0132] Sq = 7
[0133] }
[0134] else
[0135] Sq = 3
[0136] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (the strong filter condition), which is defined as follows. In condition 3, StrongFilterCondition, the following variables are derived:
[0137] Derive dpq using the HEVC method.
[0138] Derive sp3 = Abs( p3 - p0 ) using the HEVC method
[0139] if (p side is greater than or equal to 32)
[0140] if (Sp == 5)
[0141] sp3 = ( sp3 + Abs( p5 - p3 ) + 1) >> 1
[0142] else
[0143] sp3 = ( sp3 + Abs( p7 - p3 ) + 1) >> 1
[0144] Derive sq3 = Abs( q0 - q3 ) using the HEVC method.
[0145] if (q side is greater than or equal to 32)
[0146] If (Sq == 5)
[0147] sq3 = ( sq3 + Abs( q5 - q3 ) + 1) >> 1
[0148] else
[0149] sq3 = ( sq3 + Abs( q7 - q3 ) + 1) >> 1
[0150] According to HEVC, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (3)). β >> 5), and Abs(p0 - q0) is less than (5). tC + 1 ) >> 1) ? TRUE : FALSE.
[0151] 2.4.4 A more robust deblocking filter for luminance (designed for larger blocks)
[0152] When samples on either side of the boundary belong to a large block, a bilinear filter is used. Samples belonging to a large block are defined as those with a vertical edge width >= 32 and a horizontal edge height >= 32. The bilinear filter is listed below. Then, in the above HEVC deblocking, the block boundary samples pi from i=0 to Sp-1 and qi from j=0 to Sq-1 (pi and qi are the i-th sample in the row used for filtering the vertical edge, or the i-th sample in the column used for filtering the horizontal edge) are replaced by the following linear interpolation:
[0153]
[0154]
[0155] in and The item is the amplitude limit related to the above-mentioned location, and , , , and It is given below.
[0156] 2.4.5 Color deblocking control
[0157] A strong chroma filter is used on both sides of the block boundary. Here, the chroma filter is selected when both sides of the chroma edge are greater than or equal to 8 (chroma position) and the following decision with three conditions is met: The first is for the boundary strength and the decision on the size of the block. The filter can be applied when the block width or height orthogonally spanning the block edge in the chroma sample domain is equal to or greater than 8. The second and third are essentially the same as those used for HEVC luminance deblocking decisions, which are the on / off decision and the strong filter decision, respectively.
[0158] In the first decision, the boundary strength (bS) is modified for chroma filtering, and conditions are checked sequentially. If a condition is met, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS equals 2, or when bS equals 1 when a large boundary is detected. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.
[0159] In the second condition, d is then derived using the HEVC luminance deblocking method. The second condition will be true when d is less than β. In the third condition, StrongFilterCondition is derived as follows:
[0160] Derive dpq using the HEVC method;
[0161] Derive sp3 = Abs(p3 - p0) using the HEVC method; and
[0162] The formula sq3 = Abs(q0 - q3) is derived using the HEVC method.
[0163] According to the HEVC design, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (β >> 3), and Abs(p0 - q0) < (5). tC + 1 ) >> 1).
[0164] 2.4.6 Strong Deblocking Filter for Chromaticity
[0165] Strong deblocking filters for the following chroma are defined:
[0166] p2′= (3 p3+2 p2+p1+p0+q0+4) >> 3
[0167] p1′= (2 p3+p2+2 p1+p0+q0+q1+4) >> 3
[0168] p0′= (p3+p2+p1+2 p0+q0+q1+q2+4) >> 3
[0169] The example chroma filter performs deblocking on a 4×4 chroma sample grid.
[0170] 2.4.7 Location-Related Limiting
[0171] Position-dependent limiting (tcPD) is applied to the output samples of a brightness filtering process involving modifications to strong and long filters at the boundaries of 7, 5, and 3 samples. Assuming a quantization error distribution, the limiting value can be increased for samples expected to have higher quantization noise, thus anticipating a higher deviation between the reconstructed sample value and the true sample value.
[0172] For each P or Q boundary filtered by an asymmetric filter, a position-related threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) that serve as edge information, based on the results of the decision process:
[0173] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};
[0174] tcPD = (Sp == 3) ? Tc3 : Tc7;
[0175] tcQD = (Sq == 3) ? Tc3 : Tc7;
[0176] For P or Q boundaries filtered by a short symmetric filter, apply a lower-amplitude position correlation threshold:
[0177] Tc3 = { 3, 2, 1};
[0178] After defining the threshold, the filtered p'i and q'i sample values are limited according to the tcP and tcQ limiting values:
[0179] p''i = Clip3(p'i + tcPi, p'i – tcPi, p'i );
[0180] q''j = Clip3(q'j + tcQj, q'j – tcQj, q'j );
[0181] Where p'i and q'i are the filtered sample values, p''i and q''j are the output sample values after clipping, and tcPi and tcQD are the clipping thresholds derived from the VVC tc parameters, tcPD, and tcQD. The function Clip3 is the clipping function specified by VVC.
[0182] 2.4.8 Sub-block Removal and Adjustment
[0183] To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying a maximum of 5 samples on the side using sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)), as shown in the long filter's brightness control. Extending this, sub-block deblocking is adjusted such that sub-block boundaries on the 8×8 grid near the CU or implicit TU boundaries are restricted to modifying a maximum of two samples on each side.
[0184] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0185] If (block Q's mode == SUBBLOCKMODE && edge != 0) {
[0186] if (!(implicitTU && (edge == (64 / 4))))
[0187] if (edge == 2 || edge == (orthogonalLength - 2) || edge== (56 / 4) || edge == (72 / 4))
[0188] Sp = Sq = 2;
[0189] else
[0190] Sp = Sq = 3;
[0191] else
[0192] Sp = Sq = bSideQisLargeBlk ? 5:3
[0193] }
[0194] Where edge = 0 corresponds to the CU boundary, edge = 2 or orthogonalLength-2 corresponds to the sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TU is used, then implicit TU is true.
[0195] 2.5 SAO
[0196] The input to SAO is the reconstructed samples after DB. The concept of SAO is to reduce the average sample distortion of a region by first classifying region samples into multiple categories using a selected classifier, obtaining an offset for each category, and then adding the offset to each sample of that category. The region's classifier index and offset are encoded and decoded in the bitstream. In HEVC and VVC, a region (the unit of SAO parameter signaling) is defined as a CTU.
[0197] HEVC employs two SAO types that meet the requirements of low complexity. These two types are Edge Offset (EO) and Band Offset (BO), which will be discussed in more detail below. The index of the SAO type is encoded and decoded (range [0, 2]). For EO, sample classification is based on the comparison between the current sample and its neighboring samples, and this sample classification is based on the 1-D direction pattern: horizontal, vertical, 135° diagonal, and 45° diagonal.
[0198] Figure 8 shows examples of orientation patterns used for EO sample point classification. For example, four 1-D orientation patterns for EO sample point classification are shown, including horizontal (EO category=0), vertical (EO category=1), 135° diagonal (EO category=2), and 45° diagonal (EO category=3).
[0199] For a given EO category, each sample point within the CTB is classified into one of five categories. The current sample point value (labeled "c") is compared to its two nearest neighbors along the selected 1-D pattern. The classification rules for each sample point are summarized in Table I. Categories 1 and 4 are associated with local valleys and local peaks along the selected 1-D pattern, respectively. Categories 2 and 3 are associated with concave and convex angles along the selected 1-D pattern, respectively. If the current sample point does not belong to EO categories 1-4, it is classified as category 0, and SAO is not applied.
[0200] Table 3: Sampling classification rules for edge offset
[0201]
[0202] 2.6 Adaptive Loop Filter Based on Geometric Transformation in JEM
[0203] The input to DB is the reconstructed sample points after DB and SAO. The sample point classification and filtering processes are based on the reconstructed sample points after DB and SAO.
[0204] In JEM, a geometric transformation-based adaptive loop filter (GALF) with block-based filter adaptation is applied [3]. For the luminance component, one of 25 filters is selected for each 2×2 block based on the direction and activity of the local gradient.
[0205] 2.6.1 Filter Shape
[0206] Figure 9 shows the GALF filter shapes, including a 5×5 rhombus on the left, a 7×7 rhombus in the middle, and a 9×9 rhombus on the right. In JEM, up to three rhombus filter shapes can be selected for the luma component (as shown in Figure 9). The filter shape used for the luma component is indicated at the image level by a signal transmission index. Each square represents a sample point, and Ci (i = 0~6 (left), 0~12 (middle), 0~20 (right)) represents the coefficient to be applied to the sample point. For the chroma component in the image, the 5×5 rhombus shape is always used.
[0207] 2.6.1.1 Block Classification
[0208] Each 2×2 block is classified into one of 25 categories. The classification index C is based on its directionality. and activity The quantization value is derived as follows:
[0209]
[0210] In order to calculate and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using 1-D Laplacian:
[0211]
[0212] index and This refers to the coordinates of the top left sample point in the 2×2 block, and Indicator coordinates The reconstructed sample points. Then the gradients in the horizontal and vertical directions. The maximum and minimum values are set as follows:
[0213]
[0214] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0215]
[0216] In order to derive directionality The values are compared with each other and with two thresholds. and Compare:
[0217] Step 1. If and If both are true, then Set as ;
[0218] Step 2. If If yes, continue from step 3; otherwise, continue from step 4.
[0219] Step 3. If ,but Set as ,otherwise Set as ;
[0220] Step 4. If ,but Set as ,otherwise Set as .
[0221] Activity value Calculated as:
[0222]
[0223] It is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as For the two chromaticity components in the image, no classification method is applied; that is, a single set of ALF coefficients is applied for each chromaticity component.
[0224] 2.6.1.2 Geometric Transformation of Filter Coefficients
[0225] Figure 10 illustrates examples of relative coordinates used for 5×5 rhombus filter support, including diagonal, vertical flip, and rotation, respectively. Geometric transformations (such as rotation or diagonal and vertical flip) are applied to the filter coefficients associated with coordinates (k, l) before filtering each 2×2 block. This depends on the gradient values computed for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make the different blocks more similar by aligning the directions of the different blocks to which the ALF is applied.
[0226] Three geometric transformations are introduced: diagonal, vertical flip, and rotation.
[0227]
[0228] in It is the size of the filter, and These are coefficient coordinates, which make the position... In the top left corner, and in position In the bottom right corner, the transform is applied to the filter coefficients f(k, l), depending on the gradient values calculated for that block. The relationship between the transform and the four gradients in the four directions is summarized in Table 4. Figure 10 shows the transformed coefficients at each position based on the 5×5 rhombus.
[0229] Table 4. Mapping of gradients and transformations computed for a block
[0230]
[0231] 2.6.1.3 Filter Parameter Signaling
[0232] In JEM, GALF filter parameters are transmitted via signaling for the first CTU, i.e., after the stripe header and before the SAO parameters of the first CTU. Up to 25 sets of luminance filter coefficients can be transmitted via signaling. To reduce bit overhead, filter coefficients from different categories can be merged. Furthermore, the GALF coefficients of the reference image are stored and can be reused as the GALF coefficients of the current image. The current image can optionally use the GALF coefficients stored for the reference image and bypass GALF coefficient signaling. In this case, only the index of one image from the reference image is transmitted via signaling, and the stored GALF coefficients of the indicated reference image are inherited for the current image.
[0233] To support GALF temporal prediction, a candidate list of GALF filter sets is maintained. The candidate list is empty when decoding a new sequence. After decoding an image, the corresponding filter set can be added to the candidate list. Once the size of the candidate list reaches the maximum allowed value (e.g., 6 in JEM), new filter sets overwrite the oldest set in decoding order; that is, a first-in, first-out (FIFO) rule is applied to update the candidate list. To avoid duplication, a set can only be added to the list if the corresponding image does not use GALF temporal prediction. To support temporal scalability, there are multiple candidate lists of filter sets, and each candidate list is associated with a temporal layer. More specifically, each array assigned by the temporal layer index (TempIdx) can constitute a filter set for previously decoded images with a TempIdx equal to or less than TempIdx. For example, the k-th array is assigned to be associated with a TempIdx equal to k, and it contains only filter sets from images with TempIdx less than or equal to k. After a specific image is encoded or decoded, the filter set associated with that image will be used to update those arrays associated with TempIdx that are equal to or higher than TempIdx.
[0234] Temporal prediction of GALF coefficients is used for inter-frame encoded frames to minimize signaling overhead. For intra-frame frames, temporal prediction is not available, and a set of 16 fixed filters is assigned to each category. To indicate the use of fixed filters, a flag for each category is transmitted via signaling, and, if necessary, the index of the selected fixed filter is also transmitted via signaling. Even if a fixed filter is selected for a given category, the coefficients of an adaptive filter can still be sent for that category. In this case, the coefficients of the filter to be applied to the reconstructed image are the sum of two sets of coefficients.
[0235] The filtering process for the luminance component can be controlled at the CU level. A signal transmission flag indicates whether GALF is applied to the luminance component of the CU. For the chrominance component, whether GALF is applied is only indicated at the image level.
[0236] 2.6.1.4 Filtering process
[0237] On the decoder side, when GALF is enabled for a block, each sample within the block... Filtering causes sample values As shown below, where L represents the filter length, Represents the filter coefficients, and This represents the decoded filter coefficients.
[0238] (10)
[0239] Figure 11 shows an example of relative coordinates used in support of a 5×5 diamond filter, assuming the coordinates (i, j) of the current sample point are (0, 0). Sample points at different coordinates filled with the same color are multiplied by the same filter coefficient.
[0240] 2.7 Adaptive Loop Filter Based on Geometric Transformation (GALF) in VVC
[0241] 2.7.1 GALF in VTM-4
[0242] In VTM 4.0, the filtering process of the adaptive loop filter is performed as follows:
[0243] (11)
[0244] Among the sample points These are the input samples. These are the filtered output samples (i.e., the filter result), and This represents the filter coefficients. In practice, VTM 4.0 uses integer arithmetic to achieve fixed-point precision calculations:
[0245] (12)
[0246] Where L represents the filter length, and where These are the filter coefficients in fixed-point precision.
[0247] Compared to GALF in JEM, the example design of GALF in VVC has the following main changes:
[0248] 1) Adaptive filter shapes have been removed. Only 7×7 filter shapes are allowed for the luminance component, and only 5×5 filter shapes are allowed for the chrominance component.
[0249] 2) The signaling for ALF parameters has been removed from the strip / picture level to the CTU level.
[0250] 3) Class index calculation is performed at a 4×4 level instead of a 2×2 level. Additionally, as proposed in JVET-L0147, a Laplacian calculation method using downsampling for ALF classification is used. More specifically, it is unnecessary to calculate the horizontal / vertical / 45-degree diagonal / 135-degree gradient for every sample point within a block. Instead, 1:2 downsampling is used.
[0251] 2.8 Nonlinear ALF in VVC
[0252] 2.8.1 Nonlinear Filtering Reconstruction
[0253] Equation (11) can be reconstructed into the following expression without affecting encoding and decoding efficiency:
[0254] (13)
[0255] in These are the same filter coefficients as in equation (11) [except for] In equation (13), it equals 1, while in equation (11) it equals ].
[0256] Using the filter formula above (13), VVC introduces nonlinearity by using a simple limiting function to reduce the nonlinearity in the vicinity of sample values ( ) and the current sample value being filtered ( When the difference is too large, the influence of neighboring sample values is reduced, thus making ALF more efficient. More specifically, the ALF filter is modified as follows:
[0257] (14)
[0258] in It is a limiting function, and It is the limiting parameter, which depends on Filter coefficients. The encoder performs optimization to find the optimal values. .
[0259] In the JVET-N0242 implementation, a limiting parameter is specified for each ALF filter. For each filter coefficient, a limiting value is transmitted via signaling. This means that up to 12 limiting values can be transmitted via signaling in the bitstream for each luminance filter, and up to 6 limiting values can be transmitted via signaling in the bitstream for each chroma filter. To limit signaling costs and encoder complexity, only 4 fixed values are used for both inter-frame and intra-frame stripes.
[0260] Because the variance of local differences in luminance is typically higher than that in chrominance, two different sets are applied for luminance and chrominance filters. A maximum sample value is also introduced in each set (here, 1024 for a 10-bit bit depth), allowing clipping to be disabled if unnecessary.
[0261] Table 5 provides the set of limiting values used in the JVET-N0242 test. The four values were selected from the entire range of sample values for luminance (encoded and decoded on 10 bits) that are approximately equally divided in the logarithmic domain, and the range of chromaticity from 4 to 1024. More precisely, the luminance table of limiting values has been obtained using the following formula:
[0262] AlfClip L Where M=2 10 And N=4 (15)
[0263] Similarly, the colorimetric table for the limiting values is obtained according to the following formula:
[0264] AlfClip C Where M=2 10 N=4, A=4 (16)
[0265] Table 5: Authorized Amplitude Limits
[0266]
[0267] The selected limiting value is encoded in the "alf_data" syntax element using the Golomb encoding / decoding scheme corresponding to the index of the limiting values in Table 5 above. This encoding / decoding scheme is the same as that used for the filter index.
[0268] 2.9 Convolutional Neural Network-Based Loop Filters for Video Encoding and Decoding
[0269] 2.9.1 Convolutional Neural Networks
[0270] In deep learning, convolutional neural networks (CNNs or ConvNets) are a type of deep neural network most commonly used for analyzing visual images. They have been very successful in image and video recognition / processing, recommender systems, image classification, medical image analysis, and natural language processing.
[0271] CNNs are a regularized version of multilayer perceptrons. Multilayer perceptrons typically refer to fully connected networks, where every neuron in one layer is connected to all neurons in the next layer. This "full connectivity" makes them prone to overfitting data. Typical regularization methods involve adding some form of magnitude measurement of the weights to the loss function. CNNs take a different approach to regularization: they utilize hierarchical patterns in the data and assemble more complex patterns using smaller, simpler patterns. Therefore, CNNs are at the lower extremes in terms of connectivity and complexity.
[0272] Compared to other image classification / processing algorithms, CNNs use relatively little preprocessing. This means the network learns filters that are hand-designed in traditional algorithms. This independence from prior knowledge and human intervention in feature design is a major advantage.
[0273] 2.9.2 Deep Learning for Image / Video Encoding and Decoding
[0274] Deep learning-based image / video compression generally has two meanings: end-to-end compression based purely on neural networks [1, 2] and frameworks enhanced by neural networks [3, 4, 5, 6]. The first type usually adopts an autoencoder-like structure, implemented through convolutional neural networks or recurrent neural networks. Although relying purely on neural networks for image / video compression avoids any manual optimization or hand-design, the compression efficiency may be unsatisfactory. Therefore, the work in the second type uses neural networks as an aid and enhances traditional compression frameworks by replacing or enhancing some modules. In this way, they can inherit the advantages of highly optimized frameworks. For example, Li et al. proposed a fully connected network for intra-frame prediction in HEVC [3]. In addition to intra-frame prediction, deep learning has also been used to enhance other modules. For example, Dai et al. replaced the loop filter of HEVC with a convolutional neural network and achieved promising results [4]. The work in [5] applied neural networks to improve the arithmetic codec engine.
[0275] 2.9.3 Loop Filtering Based on Convolutional Neural Networks
[0276] In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, thus introducing distortion into the reconstructed frame. To mitigate this distortion, convolutional neural networks can be trained to learn the mapping from distorted frames to the original frames. In practice, this training must be performed before deploying CNN-based loop filtering.
[0277] 2.9.3.1 Training
[0278] The goal of the training process is to find the optimal values of the parameters, including the weights and biases.
[0279] First, codecs (e.g., HEVC test model (HM), JEM, VTM, etc.) are used to compress the training dataset to generate distorted reconstructed frames.
[0280] The reconstructed frames are then fed into the CNN, and the cost is calculated using the CNN's output and the ground truth frames (original frames). Common cost functions include SAD (Sum of Absolute Differences) and MSE (Mean Squared Error). Next, the gradient of the cost with respect to each parameter is derived using a backpropagation algorithm. The gradients are then used to update the parameter values. This process is repeated until the convergence criterion is met. After training is complete, the derived optimal parameters are saved for use in the inference phase.
[0281] 2.9.3.2 Convolution Process
[0282] Figure 12A shows an example CNN filter. M represents the number of feature maps. N represents the number of samples in one dimension. Figure 12B shows an example construction of the residual block (ResBlock) in Figure 12A.
[0283] During convolution, the filter moves across the image from left to right and from top to bottom, changing a column of pixels on horizontal movement and a row of pixels on vertical movement. The amount of movement the filter makes between the input images is called the stride, and it is almost always symmetrical in the height and width dimensions. The default stride, or 2D stride, is (1,1) for both height and width movements.
[0284] In most deep convolutional neural networks, residual blocks are used as basic modules and stacked multiple times to build the final network. In one example, as shown in Figure 12B, residual blocks are obtained by combining convolutional layers, modified linear unit (ReLU) / parameterized modified linear unit (PReLU) activation functions, and convolutional layers.
[0285] 2.9.3.3 Reasoning
[0286] During the inference phase, distorted reconstructed frames are fed into the CNN and processed by the CNN model whose parameters have been determined during the training phase. The input samples to the CNN can be reconstructed samples before or after DB, or before or after SAO, or before or after ALF.
[0287] 3. The technical problem solved by the disclosed technical solution
[0288] The example design of NN-based super-resolution for video encoding and decoding has the following problems.
[0289] First, the side information generated during compression can be used as additional input to improve the performance of NN-based loop filters. For example, predicted image, strip type, boundary strength, basic QP, strip QP, and intra-frame prediction, inter-frame prediction, and bidirectional inter-frame prediction (IPB) information can be used as side information.
[0290] Second, the same convolutions are used for each input in the NN-based super-resolution. However, different inputs can have different importance. Therefore, it is reasonable to assign different convolutions with different kernel sizes and number of channels to each input.
[0291] Third, it has a core size K K-type convolutions are widely used in basic residual blocks for neural network-based super-resolution. However, to design low-complexity super-resolution networks, they can be decomposed into combinations of multiple convolutions with smaller kernel sizes to reduce complexity. Furthermore, Convolution can be used together with decomposed convolutional layers.
[0292] Fourth, different neural network-based super-resolution models were used separately to generate upsampled luma and chroma components. However, a single model could be used to generate both upsampled luma and chroma components to save complexity and model storage.
[0293] Fifth, using rotation / flip operations to compress the input sequence and selecting the optimal encoding / decoding mode can improve the encoding / decoding performance of super-resolution, which is not used in current NN-based super-resolution video encoding and decoding.
[0294] 4. List of Solutions and Implementation Examples
[0295] The detailed list below should be considered as examples for explaining general concepts. These examples should not be interpreted in a narrow sense. Furthermore, these examples can be combined in any way.
[0296] One or more neural network (NN) super-resolution (SR) models are trained as part of an upsampling filtering technique used in the post-processing stage to reduce distortion generated during compression and upsample the resolution. Samples with different characteristics are processed by different NN SR models. This design details how a unified NN SR model can be designed by feeding at least one indicator as input to the NN filter, which can be related to a quality level (e.g., QP or constant rate factor (CRF) value or bit rate) / strip type / encoding / decoding mode / encoded / decoded information.
[0297] In this disclosure, the NN can be any kind of NN, such as a convolutional neural network (CNN), a fully connected neural network, a transformer, or a recurrent neural network.
[0298] In the following discussion, a video unit can be a sequence, image, strip, slice, brick, sub-image, CTU / CTB, CTU / CTB line, one or more CU / CB, one or more CTU / CTB, one or more VPDU (Virtual Pipeline Data Unit), or a sub-region within an image / strip / slice / brick. A parent video unit represents a unit larger than the video unit. Typically, a parent unit will contain multiple video units. For example, when the video unit is a CTU, the parent unit can be a strip, a CTU line, multiple CTUs, etc.
[0299] Design the edge information input of NN SR.
[0300] 1. To address problem 1, it is specified what side information will be used as additional input to the NN-based SR and how it will be processed.
[0301] a. In one example, edge information can be used as additional input to a neural network-based SR.
[0302] i. In one example, striped QP can be used as additional input to NN-based SR.
[0303] 1) In one example, the strip QP is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0304] a. In one example, in addition, striped QP is implemented via sliceQP. MAX_QP is normalized, where the value of MAX_QP can be 63.
[0305] ii. In one example, the basic QP can be used as additional input to an NN-based SR.
[0306] 1) In one example, the basic QP is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0307] a. In one example, in addition, the basic QP is derived from the baseQP. MAX_QP is normalized, where the value of MAX_QP can be 63.
[0308] iii. In one example, the prediction can be used as additional input to the NN-based SR.
[0309] iv. In one example, the strip type can be used as additional input to a NN-based SR.
[0310] 1) In one example, the strip type indicator is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0311] a. In one example, the stripe type value can also be a binary value that indicates whether the image to be filtered is an intra-frame stripe.
[0312] v. In one example, the IPB information of the video unit to be filtered can be used as additional input to the NN-based SR.
[0313] 1) In one example, the IPB information is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
[0314] a. In one example, IPB information can also be derived based on block prediction patterns.
[0315] i. In one example, if the current CU block is in inter-frame prediction mode, the value of the IPB information can be equal to A, where A is a constant value.
[0316] ii. In one example, if the current CU block is in intra-prediction mode, the value of the IPB information can be equal to B, where B is a constant value.
[0317] iii. In one example, if the current CU block is in IBC prediction mode, the value of the IPB information can be equal to C, where C is a constant value.
[0318] iv. In one example, if the current CU block is a one-way prediction block and only L0 is used for the current block, the value of the IPB information can be equal to D, where D is a constant value.
[0319] v. In one example, if the current CU block is a one-way prediction block and only L1 is used for the current block, the value of the IPB information can be equal to E, where E is a constant value.
[0320] vi. In one example, if the current CU block is a one-way prediction block and both L0 and L1 are used for the current block, the value of the IPB information can be equal to F, where F is a constant value.
[0321] vii. In one example, if the current CU block is a bidirectional prediction block, the value of the IPB information can be equal to G, where G is a constant value.
[0322] vi. In one example, the chromaticity component can be upsampled to the same size as the luminance component and used as input to a neural network-based SR.
[0323] 1) In one example, the chroma component can be upsampled by a CNN with a stride of 2.
[0324] 2) In one example, the chromaticity component can be upsampled using a non-NN method.
[0325] a. In one example, a non-NN filter could be a nearest neighbor method.
[0326] b. In one example, a non-NN filter can be a bilinear filter, a bicubic filter, or a Lanzos filter.
[0327] c. In one example, a non-NN filter could be a reference image resampling (RPR) filter.
[0328] 3) In one example, the upsampled chroma component can be concatenated with the luminance component before being fed into the convolution.
[0329] a. In one example, the reconstructed chromaticity component is upsampled and then stitched together with the reconstructed luminance component.
[0330] b. In one example, the chromaticity component of the predicted image is upsampled and then concatenated with the luminance component of the predicted image.
[0331] c. In one example, all upsampled chromaticity components are stitched together with the reconstructed and predicted luminance components.
[0332] vii. In one example, any combination of the above edge information can be used as additional input to an NN-based loop filter.
[0333] b. In one example, convolutions for each input edge information are performed separately, and then all convolution results are concatenated with the output of the convolutions for the reconstructed image.
[0334] c. In one example, the reconstructed image and edge information are concatenated and then convolution is performed.
[0335] By applying NN SR with different numbers of channels to different inputs
[0336] 2. To solve problem 2, different convolution types can be assigned to different side information inputs.
[0337] a. In one example, convolutions share the same kernel size, and different numbers of channels can be assigned to each input.
[0338] b. In one example, the number of all or some channels in the side information input is different.
[0339] c. In one example, convolutions share the same number of convolution channels, and different convolution kernel sizes are assigned to each input.
[0340] i. In one example, Kernel size can be used for convolutions of a portion of the input, and The kernel size is used for convolution of the remaining input. K represents an integer value greater than 1.
[0341] d. In one example, different numbers of convolution channels and different kernel sizes are assigned to each input.
[0342] Simplifying the NN filter through decomposition
[0343] 3. In order to solve problem 3, Decomposition and convolution Convolutions can be combined within a basic residual block.
[0344] a. In one example, A convolution can be decomposed into a combination of multiple convolutions with smaller kernel sizes. K represents an integer value greater than 1. and These represent the number of input channels and the number of output channels of the convolution, respectively.
[0345] i. In one example, when a convolution is decomposed, the number of its input channels and output channels may remain unchanged.
[0346] 1) In one example, Convolution is decomposed into Convolution and subsequent Combinations of convolutions.
[0347] 2) In one example, Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer can be placed after each convolution.
[0348] 3) In one example, Convolution is decomposed into Convolution and subsequent Combinations of convolutions.
[0349] 4) In one example, Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer can be placed after each convolution.
[0350] ii. In one example, the number of input channels and the number of output channels can be changed when the convolution is decomposed.
[0351] 1) In one example, Convolution is decomposed into Convolution and subsequent Combinations of convolutions, where It is different Positive integers.
[0352] 2) In one example, Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer can be placed after each convolution, where It is different Positive integers.
[0353] 3) In one example, Convolution is decomposed into Convolution and subsequent Combinations of convolutions, where It is different Positive integers.
[0354] 4) In one example, Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer can be placed after each convolution, where It is different Positive integers.
[0355] iii. In one example, part The convolution is decomposed.
[0356] iv. In one example, all The convolution is decomposed.
[0357] b. In one example, Convolutional layers can be decomposed It is used before or after the convolutional layer.
[0358] i. In one example, two Convolutional layers are designed on decomposed Before the convolutional layer.
[0359] 1) In one example, furthermore, any number of activation layers can be placed in any Convolutional layers may not be placed after any After the convolutional layer.
[0360] 2) In one example, the first Convolution has the number of input channels and number of output channels .second Convolution has the number of input channels and number of output channels .
[0361] 3) In one example, the following constraints also apply: .
[0362] ii. In one example, any number Convolutional layers are designed on decomposed Before the convolutional layer, and it can be viewed as multiple sets of two... Combination of convolutional layers.
[0363] iii. In one example, a single Convolutional layers are designed on decomposed After the convolutional layer.
[0364] 1) In one example, in addition, the activation layer can be placed It may not be necessary to place it after the convolutional layer. After the convolutional layer.
[0365] iv. In one example, Convolutional layers can simultaneously decompose It is used before and after the convolutional layer.
[0366] Regarding the output of NN SR
[0367] 4. To address problem 4, a single NN-based SR can be used to generate the output of the luminance component and the output of the chrominance component.
[0368] a. In one example, two branches can be designed in a single neural network such that one branch generates the output of the luminance component and the other branch generates the output of the chrominance component.
[0369] i. In one example, the two branches share the same input.
[0370] ii. In one example, the inputs to the two branches come from different channels of the same feature map.
[0371] iii. In one example, the two branches can be designed to have the same network structure.
[0372] 1) In one example, each branch also consists of multiple basic blocks. A basic block is any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU).
[0373] a. In one example, the basic block may or may not use a residual structure.
[0374] iv. In one example, the two branches can be designed with different network structures.
[0375] 1) In one example, each branch also consists of multiple basic blocks. A basic block is any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU) or other layers used in the neural network.
[0376] a. In one example, the basic block may or may not use a residual structure.
[0377] 2) In one example, the branch that generates the chroma component can use fewer convolutions / channels / basic blocks compared to the branch that generates the luminance component.
[0378] v. In one example, an upsampling layer can be designed in the branch of the luminance component or the branch of the chrominance component, or in both branches.
[0379] 1) In one example, the pixel rearrangement method is used as an upsampling layer.
[0380] 2) In one example, a transposed convolution with a stride of 2 is used as an upsampling layer.
[0381] 3) In one example, the convolutional layer is designed to follow the upsampling layer.
[0382] b. In one example, a single branch can be designed in a neural network such that it together generates the luminance and chrominance components.
[0383] i. In one example, the branch also consists of multiple basic blocks. A basic block is any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU).
[0384] 1) In one example, the basic block may or may not use a residual structure.
[0385] ii. In one example, the upsampling layer can be designed for the luminance component output, the chrominance component output, or both outputs.
[0386] 1) In one example, the pixel rearrangement method is used as an upsampling layer.
[0387] 2) In one example, a transposed convolution with a stride of 2 is used as an upsampling layer.
[0388] 3) In one example, the convolutional layer is designed to follow the upsampling layer.
[0389] c. In one example, the output luminance component and output chrominance component generated by a single NN-based SR can also be used separately.
[0390] i. In one example, in addition, the output luminance component generated by a single NN-based SR can be used, and the output luminance component generated by a single NN-based SR can be left unused.
[0391] ii. In one example, in addition, the output luminance component generated by a single NN-based SR may not be used, and the output luminance component generated by a single NN-based SR may be used.
[0392] Image enhanced for optimal compression.
[0393] 5. To address problem 5, the input image can be enhanced, and the best encoding / decoding method can be selected.
[0394] a. In one example, the enhancement method is specified as follows:
[0395] i. In one example, the enhancement method is rotation at an arbitrary angle.
[0396] 1) In one example, in addition, a 90-degree rotation is applied.
[0397] 2) In one example, in addition, a 180-degree rotation is applied.
[0398] 3) In one example, a 280-degree rotation was also applied.
[0399] ii. In one example, the enhancement method is flipping.
[0400] 1) In one example, in addition, a vertical flip is applied.
[0401] 2) In one example, in addition, a horizontal flip is applied.
[0402] iii. In one example, the enhancement method can be one or more combinations of rotations and flips at arbitrary angles.
[0403] b. In one example, the best encoding / decoding method was selected based on multi-pass compression.
[0404] i. In one example, the original input image and all enhanced input images are compressed, and the best encoding / decoding method is selected based on the rate-distortion optimization (RDO) choice of these compression results.
[0405] c. In one example, after the best encoding / decoding method is selected, the enhancement method is transmitted via signaling.
[0406] d. In one example, the decoder can generate an upsampled reconstruction based on the inverse operation of the enhancement method transmitted through the signal.
[0407] Design NN SR
[0408] 6. It is proposed to design NN SR by using all or part of the items mentioned, which should not be interpreted in a narrow sense.
[0409] 5. Examples
[0410] 5.1 Example 1
[0411] Figure 13A shows an example network structure for NN SR, which includes all items 1 through 4. Figure 13B shows an example backbone block used in Figure 13A. The hyperparameters of the SR network structure are specified in the table below.
[0412]
[0413] 5.2 Example 2
[0414] Figure 14A shows an example network structure for NN SR, which includes all items 1 through 4. Figure 14B shows the example backbone block used in Figure 14A. The hyperparameters of the SR network structure are specified in the table below.
[0415]
[0416] 6. References
[0417] [1] Johannes Ballé, Valero Laparra, and Eero P Simoncelli. 2016. End-to-end optimization of nonlinear transform codes for perceptual quality. InPCS. IEEE, 1–5.
[0418] [2] Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszár.2017. Lossy image compression with compressive autoencoders. arXiv preprintarXiv:1703.00395 (2017).
[0419] [3] Jiahao Li, Bin Li, Jizheng Xu, Ruiqin Xiong, and Wen Gao. 2018.Fully Connected Network-Based Intra Prediction for Image Coding. IEEETransactions on Image Processing 27, 7 (2018), 3236–3247.
[0420] [4] Yuanying Dai, Dong Liu, and Feng Wu. 2017. A convolutional neuralnetwork approach for post-processing in HEVC intra coding. In MMM. Springer,28–39.
[0421] [5] Rui Song, Dong Liu, Houqiang Li, and Feng Wu. 2017. Neuralnetwork-based arithmetic coding of intra prediction modes in HEVC. In VCIP.IEEE, 1–4.
[0422] [6] J Pfaff, P Helle, D Maniry, S Kaltenstadler, W Samek, H Schwarz, D Marpe, and T Wiegand. 2018. Neural network based intra prediction for videocoding. In Applications of Digital Image Processing XLI, Vol. 10752. International Society for Optics and Photonics, 1075213.
[0423] Figure 15 is a block diagram illustrating an example video processing system 4000 in which various embodiments disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0424] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this disclosure. Encoding / decoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding / decoding component 4004 to produce an encoded / decoded representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or encoded / decoded) representation of the video received at input 4002, whether stored or communicated, may be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding / decoding tools or operations are used by the encoder, and the corresponding decoding tools or operations that reverse the encoding / decoding results are performed by the decoder.
[0425] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The technologies described in this document can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0426] Figure 16 is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors 4102(s) may be configured to implement one or more methods described herein. The memories 4104(s) may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, for example, a graphics coprocessor.
[0427] Figure 17 is a flowchart of an example method 4200 for video processing. Method 4200 includes determining, in step 4202, an applied stripe quantization parameter (QP) as additional input to a neural network (NN)-based super-resolution (SR) process. In step 4204, a conversion between visual media data and a bitstream is performed based on the NN-based SR. According to the example, the conversion in step 4204 may include encoding at an encoder or decoding at a decoder.
[0428] It should be noted that method 4200 can be implemented in a means of processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to execute method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video encoding / decoding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video encoding / decoding device executes method 4200.
[0429] Figure 18 is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize the techniques of this disclosure. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.
[0430] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0431] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.
[0432] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVM) standard, and other existing and / or further standards.
[0433] Figure 19 is a block diagram illustrating an example of a video encoder 4400, which may be the video encoder 4314 in the system 4300 shown in Figure 18. The video encoder 4400 may be configured to perform any or all of the techniques of this disclosure. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0434] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.
[0435] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0436] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.
[0437] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0438] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).
[0439] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0440] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0441] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0442] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images of list 0, and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0443] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0444] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0445] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0446] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0447] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0448] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0449] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0450] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0451] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0452] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.
[0453] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce artifacts in the video block.
[0454] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0455] Figure 20 is a block diagram illustrating an example of a video decoder 4500, which may be the video decoder 4324 in the system 4300 shown in Figure 18. The video decoder 4500 may be configured to perform any or all of the techniques of this disclosure. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0456] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.
[0457] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine such information, for example, by executing AMVP and Merge modes.
[0458] The motion compensation unit 4502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0459] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.
[0460] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0461] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.
[0462] The reconstruction unit 4506 can sum the residual block with the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0463] Figure 21 is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and by utilizing the side information from the encoding and decoding to reduce the mean square error between the original and reconstructed samples through signal transmission offset and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be considered as a tool to attempt to capture and repair artifacts caused by previous stages.
[0464] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy encoding / decoding component 4618. The entropy encoding / decoding component 4618 entropy-encodes and decodes the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0465] The following is a list of some preferred solutions.
[0466] The following solutions illustrate examples of the techniques discussed in this article.
[0467] 1. A method for processing video data, comprising: determining to apply side information as additional input to a neural network (NN)-based super-resolution (SR); and performing a conversion between visual media data and a bitstream based on the NN-based SR.
[0468] 2. The method according to Solution 1, wherein: the strip QP is used as additional input to the NN-based SR, or the strip QP is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered, or the strip QP is processed by sliceQP. MAX_QP is normalized, and the value of MAX_QP is 63.
[0469] 3. The method according to any one of solutions 1-2, wherein: the base QP is used as additional input to the NN-based SR, or the base QP is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered, or the base QP is obtained through a baseQP. MAX_QP is normalized, where the value of MAX_QP is 63.
[0470] 4. The method according to any one of solutions 1-3, wherein the prediction is used as an additional input to the NN-based SR.
[0471] 5. The method according to any one of solutions 1-4, wherein: the strip type is used as additional input to the NN-based SR, or the strip type indicator is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered, or the strip type value is a binary value indicating whether the image to be filtered is an intra-frame stripe.
[0472] 6. The method according to any one of solutions 1-5, wherein: the IPB information of the video unit to be filtered is used as additional input to the NN-based SR, or the IPB information is first sliced or expanded into a two-dimensional array of the same size as the video unit to be filtered, or the IPB information is derived based on the block prediction mode, or if the current CU block is in inter-frame prediction mode, the value of the IPB information is equal to A, where A is a constant value; or if the current CU block is in intra-frame prediction mode, the value of the IPB information is equal to B, where B is a constant value; or if the current CU block is in IBC prediction mode. If the current CU block is a one-way prediction block and only L0 is used for the current block, then the value of the IPB information is equal to C, where C is a constant value; or if the current CU block is a one-way prediction block and only L1 is used for the current block, then the value of the IPB information is equal to E, where E is a constant value; or if the current CU block is a one-way prediction block and both L0 and L1 are used for the current block, then the value of the IPB information is equal to F, where F is a constant value; or if the current CU block is a two-way prediction block, then the value of the IPB information is equal to G, where G is a constant value.
[0473] 7. The method according to any one of solutions 1-6, wherein: the chroma component is upsampled to the same size as the luminance component as input to a neural network-based SR, or the chroma component is upsampled by a CNN with a stride of 2, or the chroma component is upsampled by a non-NN filter, or the non-NN filter is a nearest neighbor filter, or the non-NN filter is a bilinear filter, a bicubic filter, or a Lanzos filter, or the non-NN filter is an RPR filter, or the upsampled chroma component is concatenated with the luminance component before being fed into a convolution, or the reconstructed chroma component is upsampled and then concatenated with the reconstructed luminance component, or the chroma component of the predicted image is upsampled and then concatenated with the luminance component of the predicted image, or all upsampled chroma components are concatenated together with the reconstructed and predicted luminance components.
[0474] 8. The method according to any one of solutions 1-7, wherein: convolution for each input side information is performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image, or the reconstructed image and the side information are concatenated and then convolution is performed.
[0475] 9. The method according to any one of solutions 1-8, wherein different convolution types are assigned to different side information inputs.
[0476] 10. The method according to any one of solutions 1-9, wherein: the convolutions share the same convolution kernel size and different numbers of channels are assigned to each input, or all or some of the channel numbers of the side information inputs are different, or the convolutions share the same number of convolution channels and different convolution kernel sizes are assigned to each input, or The kernel size was used for the convolution of a portion of the input, and The kernel size is used for the convolution of the remaining inputs. K represents an integer value greater than 1, or different numbers of convolution channels and different kernel sizes are assigned to each input.
[0477] 11. The method according to any one of solutions 1-10, wherein, Decomposition and convolution The use of convolution is combined within a basic residual block.
[0478] 12. The method according to any one of solutions 1-11, wherein: Convolution is decomposed into a combination of multiple convolutions with smaller kernel sizes, where K represents an integer value greater than 1. and These represent the number of input channels and the number of output channels of the convolution, respectively. Alternatively, when the convolution is decomposed, the number of input channels and the number of output channels of the convolution remain unchanged. Convolution is decomposed into Convolution and subsequent Combinations of convolutions, or as described Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer is placed after each convolution, or as described above. Convolution is decomposed into Convolution and subsequent Combinations of convolutions, or as described Convolution is decomposed into Convolution and subsequent The convolutions are combined, and any activation layer is placed after each convolution, or the number of input and output channels of the convolutions changes when the convolutions are decomposed, or the... Convolution is decomposed into Convolution and subsequent Combinations of convolutions, where It is different A positive integer, or the aforementioned Convolution is decomposed into Convolution and subsequent A combination of convolutions, with any activation layer placed after each convolution, where, It is different A positive integer, or the aforementioned Convolution is decomposed into Convolution and subsequent Combinations of convolutions, where It is different A positive integer, or the aforementioned Convolution is decomposed into Convolution and subsequent A combination of convolutions, with any activation layer placed after each convolution, where It is different Positive integers, or parts Convolution is decomposed, or all of them are decomposed. The convolution is decomposed.
[0479] 13. The method according to any one of solutions 1-12, wherein: Convolutional layers after decomposition It is used before or after the convolutional layer, or both. Convolutional layers are designed on decomposed Before a convolutional layer, or any number of activation layers, can be placed in any... Convolutional layers may not be placed after any After the convolutional layer, or the first Convolution has the number of input channels and number of output channels And the second Convolution has the number of input channels and number of output channels Or apply the following constraints: and or any number Convolutional layers are designed on decomposed Before the convolutional layer, and it is considered as multiple sets of two... Combinations of convolutional layers, or single convolutional layers Convolutional layers are designed on decomposed After the convolutional layer, or the activation layer, it can be placed... It may not be necessary to place it after the convolutional layer. After the convolutional layer, or Convolutional layers are simultaneously decomposed It is used before and after the convolutional layer.
[0480] 14. The method according to any one of solutions 1-13, wherein a single NN-based SR is used to generate the output of the luminance component and the output of the chrominance component.
[0481] 15. The method according to any one of solutions 1-14, wherein: two branches are designed in the single neural network such that one branch generates the output of the luminance component and the other branch generates the output of the chrominance component, or the two branches share the same input, or the inputs of the two branches come from different channels of the same feature map, or the two branches are designed to have the same network structure, or each branch consists of multiple basic blocks. The basic block is any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU), or the basic block may or may not use a residual structure, or the two branches are designed with different network structures, or each branch consists of multiple basic blocks, and the basic block is any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU) or other layers used in the neural network, or the basic block may or may not use a residual structure, or the branch generating the chroma component uses fewer convolutions / channels / basic blocks compared to the branch generating the luminance component, or an upsampling layer is designed in the branch generating the luminance component or the branch generating the chroma component, or an upsampling layer is designed in both branches, or a pixel rearrangement method is used as the upsampling layer, or a transposed convolution with a stride of 2 is used as the upsampling layer, or a convolutional layer is designed after the upsampling layer.
[0482] 16. The method according to any one of solutions 1-15, wherein: a single branch in the neural network is designed such that it generates the luminance component and the chrominance component together, or the branch consists of multiple basic blocks, the basic blocks being any combination of convolutional layers / fully connected layers / transformer layers / activation layers (such as ReLU or PReLU), or the basic blocks may or may not use residual structures, or the upsampling layer is designed for the luminance component output, the chrominance component output, or the two outputs, or a pixel rearrangement method is used as the upsampling layer, or a transposed convolution with a stride of 2 is used as the upsampling layer, or a convolutional layer is designed after the upsampling layer.
[0483] 17. The method according to any one of solutions 1-16, wherein: the output luminance component and the output chrominance component generated by the single NN-based SR are used individually, or the output luminance component generated by the single NN-based SR is used and the output luminance component generated by the single NN-based SR may not be used, or the output luminance component generated by the single NN-based SR may not be used and the output luminance component generated by the single NN-based SR is used.
[0484] 18. The method according to any one of solutions 1-17, wherein the input image is enhanced and the optimal encoding / decoding method is selected.
[0485] 19. The method according to any one of solutions 1-18, wherein: the enhancement method is specified, or the enhancement method is a rotation at any angle, or a rotation of 90 degrees is applied, or a rotation of 180 degrees is applied, or a rotation of 280 degrees is applied, or the enhancement method is a flip, or a flip along the vertical direction is applied, or a flip along the horizontal direction is applied, or the enhancement method is one or more combinations of rotation and flip at any angle.
[0486] 20. The method according to any one of solutions 1-19, wherein: the optimal encoding / decoding method is selected based on multi-pass compression, or the original input and all enhanced input images are compressed, and the optimal encoding / decoding method is selected based on the RDO selection of these compression results.
[0487] 21. The method according to any one of solutions 1-20, wherein: after the optimal encoding / decoding method is selected, the enhancement method is transmitted via signaling, or the decoder can generate the upsampled reconstruction based on the inverse operation of the enhancement method transmitted via signaling.
[0488] 22. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-21.
[0489] 23. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method according to any one of solutions 1-21.
[0490] 24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining application side information as additional input to a neural network (NN)-based super-resolution (SR); and generating the bitstream based on the determination.
[0491] 25. A method for storing a bitstream of video, comprising: determining application side information as additional input to a neural network (NN)-based super-resolution (SR); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0492] 26. A method, apparatus, or system described in this document.
[0493] The following solutions illustrate other examples of the techniques discussed in this article.
[0494] 1. A method for processing video data, comprising: determining an application strip quantization parameter (QP) as additional input to a neural network (NN)-based super-resolution (SR) process; and performing a conversion between visual media data and a bitstream based on the NN-based SR process.
[0495] 2. The method according to Solution 1, wherein the strip QP is first sliced or expanded into a two-dimensional (2D) array with the same size as the video unit to be filtered before being applied as input to the NN-SR process.
[0496] 3. The method according to any one of solutions 1-2, wherein the strip QP is transmitted via sliceQP. MAX_QP is normalized, where the value of MAX_QP is 63.
[0497] 4. The method according to any one of solutions 1-3, wherein the basic QP is used as an additional input to the NN-based SR process.
[0498] 5. The method according to any one of solutions 1-4, wherein the basic QP is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered before being applied as input to the NN-SR process.
[0499] 6. The method according to any one of solutions 1-5, wherein the basic QP is derived from baseQP. MAX_QP is normalized, where the value of MAX_QP is 63.
[0500] 7. The method according to any one of solutions 1-6, wherein the prediction is used as an additional input to the NN-based SR.
[0501] 8. The method according to any one of solutions 1-7, wherein the strip type is used as an additional input to the NN-based SR.
[0502] 9. The method according to any one of solutions 1-8, wherein the inter-frame prediction mode (IPB) information of the video unit to be filtered is used as an additional input to the NN-based SR.
[0503] 10. The method according to any one of solutions 1-9, wherein the IPB information is derived based on the block prediction pattern.
[0504] 11. The method according to any one of solutions 1-10, wherein the chromaticity component is upsampled to the same size as the luminance component and used as input to the NN-based SR process.
[0505] 12. The method according to any one of solutions 1-11, wherein the chroma components are upsampled by a convolutional neural network (CNN) with a stride of 2.
[0506] 13. The method according to any one of solutions 1-12, wherein the chroma component is upsampled by a non-NN filter, wherein the non-NN filter is a reference image resampling (RPR) filter.
[0507] 14. The method according to any one of solutions 1-13, wherein the non-NN filter is a nearest neighbor method.
[0508] 15. The method according to any one of solutions 1-14, wherein the upsampled chromaticity component is concatenated with the luminance component before applying convolution.
[0509] 16. The method according to any one of solutions 1-15, wherein the reconstructed chromaticity component is upsampled and then stitched together with the reconstructed luminance component.
[0510] 17. The method according to any one of solutions 1-16, wherein all upsampled chromaticity components are stitched together with the reconstructed and predicted luminance components.
[0511] 18. The method according to any one of solutions 1-17, wherein the information is used as additional input to the NN-based loop filter.
[0512] 19. The method according to any one of solutions 1-18, wherein convolution for each input edge information is performed separately, and then all convolution results are concatenated with the output of the convolution of the reconstructed image.
[0513] 20. The method according to any one of solutions 1-19, wherein different convolution types are assigned to different side information inputs.
[0514] 21. The method according to any one of solutions 1-20, wherein different numbers of convolution channels and different kernel sizes are assigned to each input.
[0515] 22. The method according to any one of solutions 1-21, wherein, Decomposition and convolution The use of convolution is combined within a basic residual block.
[0516] 23. The method according to any one of solutions 1-22, wherein, The convolution is decomposed into a combination of multiple convolutions with smaller kernel sizes, where K represents an integer value greater than 1, and the... and stated These represent the number of input channels and the number of output channels of the convolution, respectively.
[0517] 24. The method according to any one of solutions 1-23, wherein when the convolution is decomposed, the number of input channels and the number of output channels of the convolution do not change.
[0518] 25. The method according to any one of solutions 1-24, wherein, Convolution is decomposed into Convolution and subsequent Combinations of convolutions.
[0519] 26. The method according to any one of solutions 1-25, wherein the number of input channels and the number of output channels of the convolution are changed when the convolution is decomposed.
[0520] 27. The method according to any one of solutions 1-26, wherein, Convolution is decomposed into Convolution and subsequent Combinations of convolutions, wherein, It is different from what is described Positive integers.
[0521] 28. The method according to any one of solutions 1-27, wherein: part Convolution is decomposed, or all of them are decomposed. The convolution is decomposed.
[0522] 29. The method according to any one of solutions 1-28, wherein: Convolutional layers after decomposition It is used before or after the convolutional layer, or both. Convolutional layers are designed on decomposed Before the convolutional layer.
[0523] 30. The method according to any one of solutions 1-29, wherein: an arbitrary number of activation layers are placed in an arbitrary After the convolutional layer, or the first Convolution has the number of input channels and number of output channels And the second Convolution has the number of input channels and number of output channels ,or .
[0524] 31. The method according to any one of solutions 1-30, wherein: a single Convolutional layers are designed on decomposed After the convolutional layer, or when the activation layer is placed in a single... After the convolutional layer.
[0525] 32. The method according to any one of solutions 1-31, wherein a single NN-based SR is used to generate the output of the luminance component and the output of the chrominance component.
[0526] 33. The method according to any one of solutions 1-32, wherein two branches are designed in a single neural network such that the first branch generates the output of the luminance component and the second branch generates the output of the chrominance component.
[0527] 34. The method according to any one of solutions 1-33, wherein the inputs of the two branches come from different channels of the same feature map.
[0528] 35. The method according to any one of solutions 1-34, wherein the two branches are designed with different network structures.
[0529] 36. The method according to any one of solutions 1-35, wherein: each branch comprises a plurality of basic blocks, and the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, modified linear units (ReLU), parameterized modified linear units (PReLU), or other layers used in a neural network, or the basic blocks use residual structures.
[0530] 37. The method according to any one of solutions 1-36, wherein the branch for generating the chroma component uses fewer convolutions, channels, or basic blocks compared to the branch for generating the luminance component.
[0531] 38. The method according to any one of solutions 1-37, wherein an upsampling layer is designed in the branch of the luminance component or the branch of the chrominance component, or an upsampling layer is designed in both branches.
[0532] 39. The method according to any one of solutions 1-38, wherein pixel rearrangement is used as an upsampling layer.
[0533] 40. The method according to any one of solutions 1-39, wherein: the two branches share the same input, or the two branches are designed to have the same network structure, or each branch consists of multiple basic blocks, and the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, ReLU, PReLU, or other layers used in the neural network, or the basic blocks may use residual structures or may not use residual structures, or a transposed convolution with a stride of 2 is used as an upsampling layer, or a convolutional layer is designed to follow the upsampling layer.
[0534] 41. The method according to any one of solutions 1-40, wherein the reconstructed image and edge information are stitched together and then convolved.
[0535] 42. The method according to any one of solutions 1-41, wherein: the non-NN filter is a bilinear filter, a bicubic filter, or a Lanzos filter, or the upsampled chromaticity component of the predicted image is concatenated with the luminance component of the predicted image.
[0536] 43. The method according to any one of solutions 1-42, wherein: the stripe type indicator is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered, or the stripe type value is a binary value indicating whether the picture to be filtered includes intra-frame stripes.
[0537] 44. The method according to any one of solutions 1-43, wherein: the IPB information is first sliced or expanded into a two-dimensional array having the same size as the video unit to be filtered, or when the current CU block is in inter-frame prediction mode, the value of the IPB information is equal to A, where A is a constant value; or when the current CU block is in intra-frame prediction mode, the value of the IPB information is equal to B, where B is a constant value; or when the current CU block is in IBC prediction mode, the value of the IPB information is equal to C, where C is a constant value; or when the current CU block is in IBC prediction mode... When the current CU block is a one-way prediction block and only L0 is used for the current block, the value of the IPB information is equal to D, where D is a constant value; or when the current CU block is a one-way prediction block and only L1 is used for the current block, the value of the IPB information is equal to E, where E is a constant value; or when the current CU block is a one-way prediction block and both L0 and L1 are used for the current block, the value of the IPB information is equal to F, where F is a constant value; or when the current CU block is a two-way prediction block, the value of the IPB information is equal to G, where G is a constant value.
[0538] 45. The method according to any one of solutions 1-44, wherein: convolutions share the same kernel size and different numbers of channels are assigned to each input, or the number of channels for side information inputs are different, or convolutions share the same number of convolution channels and different kernel sizes are assigned to each input, or The kernel size is used for convolution of part of the input, and The kernel size is used for convolution of the remaining input, and K represents an integer value greater than 1.
[0539] 46. The method according to any one of solutions 1-45, wherein: Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer is placed after each convolution, or Convolution is decomposed into Convolution and subsequent Combinations of convolutions, or Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer is placed after each convolution, or Convolution is decomposed into Convolution and subsequent A combination of convolutions, with any activation layer placed after each convolution, where It is different positive integers, or Convolution is decomposed into Convolution and subsequent Combinations of convolutions, wherein, It is different from what is described positive integers, or Convolution is decomposed into Convolution and subsequent The convolutions are combined, and any activation layers are placed after each convolution, wherein... It is different from what is described Positive integers.
[0540] 47. The method according to any one of solutions 1-46, wherein: any number of Convolutional layers are designed on decomposed Before the convolutional layer, and the Convolutional layers are treated as multiple groups of two Combinations of convolutional layers, or as described Convolutional layers are simultaneously decomposed It is used before and after the convolutional layer.
[0541] 48. The method according to any one of solutions 1-47, wherein: a single branch in the neural network is designed such that it together generates a luminance component and a chrominance component, or the branch comprises a plurality of basic blocks, wherein the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, ReLU, PReLU, or other layers used in the neural network, or the basic blocks use a residual structure, or an upsampling layer is designed for luminance component output, chrominance component output, or both, or pixel rearrangement is used as an upsampling layer, or a transposed convolution with a stride of 2 is used as the upsampling layer, or a convolutional layer is designed to follow the upsampling layer.
[0542] 49. The method according to any one of solutions 1-48, wherein: the output luminance component and the output chrominance component generated by the single NN-based SR are used individually, or the output luminance component generated by the single NN-based SR is used and the output luminance component generated by the single NN-based SR may not be used, or the output luminance component generated by the single NN-based SR may not be used and the output luminance component generated by the single NN-based SR is used.
[0543] 50. The method according to any one of solutions 1-49, wherein the input image is enhanced and the optimal encoding / decoding method is selected.
[0544] 51. The method according to any one of solutions 1-50, wherein: an enhancement is specified, or the enhancement is a rotation of any angle, or a rotation of 90 degrees is applied, or a rotation of 180 degrees is applied, or a rotation of 280 degrees is applied, or the enhancement is a flip, or a flip along the vertical direction is applied, or a flip along the horizontal direction is applied, or the enhancement is one or more combinations of rotation and flip at any angle.
[0545] 52. The method according to any one of solutions 1-51, wherein: the optimal encoding / decoding method is selected based on multi-pass compression, or the original input and all enhanced input images are compressed, and the optimal encoding / decoding method is selected based on rate-distortion optimization (RDO) selection of these compression results.
[0546] 53. The method according to any one of solutions 1-52, wherein: after the optimal encoding / decoding method is selected, the enhancement is transmitted via signal transmission, or the decoder generates an upsampled reconstruction based on the inverse operation of the enhancement transmitted via signal transmission.
[0547] 54. The method according to any one of solutions 1-53, wherein the conversion includes encoding the visual media data into the bitstream.
[0548] 55. The method according to any one of solutions 1-53, wherein the conversion includes decoding the visual media data from the bitstream.
[0549] 56. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-55.
[0550] 57. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method according to any one of solutions 1-55.
[0551] 58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining application side information as additional input to a neural network (NN)-based super-resolution (SR); and generating the bitstream based on the determination.
[0552] 59. A method for storing a bitstream of video, comprising: determining application side information as additional input to a neural network (NN)-based super-resolution (SR); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0553] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.
[0554] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at the same position in the bitstream defined by the syntax or bits propagated at different positions. For example, a macroblock can be encoded based on the error residual value after transformation and encoding / decoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing whether some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0555] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.
[0556] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0557] The processing and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0558] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0559] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0560] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the specific order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.
[0561] Only a few implementations and examples are described, and other implementations, improvements and variations can be made based on what is described and shown in this patent document.
[0562] When there is no intermediate component (other than a line, trace, or other medium between the first and second components), the first component is directly coupled to the second component. When there is an intermediate component between the first and second components other than a line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0563] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The examples presented are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0564] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or other). Other examples of variations, substitutions, and modifications can be determined by those skilled in the art and can be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing video data, comprising: Determine the application of strip quantization parameters (QP) as additional input to the neural network (NN) based super-resolution (SR) process; And perform the conversion between visual media data and bitstream based on the NN-based SR process.
2. The method according to claim 1, wherein, Before being applied as input to the NN-SR process, the strip QP is first sliced or expanded into a two-dimensional (2D) array with the same size as the video unit to be filtered.
3. The method according to any one of claims 1-2, wherein, The strip QP is obtained through sliceQP. MAX_QP is normalized, where the value of MAX_QP is 63.
4. The method according to any one of claims 1-3, wherein, The basic QP is used as an additional input to the NN-based SR process.
5. The method according to any one of claims 1-4, wherein, Before being applied as input to the NN-SR process, the basic QP is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered.
6. The method according to any one of claims 1-5, wherein, The basic QP is obtained through baseQP. MAX_QP is normalized, where the value of MAX_QP is 63.
7. The method according to any one of claims 1-6, wherein, The prediction is used as an additional input to the NN-based SR.
8. The method according to any one of claims 1-7, wherein, The strip type is used as an additional input to the NN-based SR.
9. The method according to any one of claims 1-8, wherein, The inter-frame prediction mode (IPB) information of the video unit to be filtered is used as additional input to the NN-based SR.
10. The method according to any one of claims 1-9, wherein, The IPB information is derived based on the block prediction pattern.
11. The method according to any one of claims 1-10, wherein, The chromaticity component is upsampled to the same size as the luminance component and used as input to the NN-based SR process.
12. The method according to any one of claims 1-11, wherein, The chromaticity components are upsampled by a convolutional neural network (CNN) with a stride of 2.
13. The method according to any one of claims 1-12, wherein, The chroma components are upsampled by non-NN filters, wherein the non-NN filter is a reference image resampling (RPR) filter.
14. The method according to any one of claims 1-13, wherein, Non-NN filters are nearest neighbor methods.
15. The method according to any one of claims 1-14, wherein, The upsampled chromaticity component is concatenated with the luminance component before applying convolution.
16. The method according to any one of claims 1-15, wherein, The reconstructed chromaticity components are upsampled and then stitched together with the reconstructed luminance components.
17. The method according to any one of claims 1-16, wherein, All upsampled chroma components are reconstructed and predicted together with the luminance components.
18. The method according to any one of claims 1-17, wherein, The information is used as additional input to the NN-based loop filter.
19. The method according to any one of claims 1-18, wherein, Convolutions for each input edge information are performed separately, and then all convolution results are concatenated with the output of the convolutions in the reconstructed image.
20. The method according to any one of claims 1-19, wherein, Different convolution types are assigned to different side information inputs.
21. The method according to any one of claims 1-20, wherein, Different numbers of convolution channels and different kernel sizes are assigned to each input.
22. The method according to any one of claims 1-21, wherein, Decomposition and convolution The use of convolution is combined within a basic residual block.
23. The method according to any one of claims 1-22, wherein, The convolution is decomposed into a combination of multiple convolutions with smaller kernel sizes, where K represents an integer value greater than 1, and the... Japanese These represent the number of input channels and the number of output channels of the convolution, respectively.
24. The method according to any one of claims 1-23, wherein, When a convolution is decomposed, the number of input channels and the number of output channels of the convolution remain unchanged.
25. The method according to any one of claims 1-24, wherein, Convolution is decomposed into Convolution and subsequent Combinations of convolutions.
26. The method according to any one of claims 1-25, wherein, When a convolution is decomposed, the number of input channels and the number of output channels of the convolution change.
27. The method according to any one of claims 1-26, wherein, Convolution is decomposed into Convolution and subsequent Combinations of convolutions, wherein, It is different from what is described Positive integers.
28. The method according to any one of claims 1-27, wherein: part The convolution is decomposed, or all of the above. The convolution is decomposed.
29. The method according to any one of claims 1-28, wherein: Convolutional layers after decomposition It is used before or after the convolutional layer, or both. Convolutional layers are designed on decomposed Before the convolutional layer.
30. The method according to any one of claims 1-29, wherein: Any number of activation layers are placed in any After the convolutional layer, or the first Convolution has the number of input channels and number of output channels And the second Convolution has the number of input channels and number of output channels ,or and 。 31. The method according to any one of claims 1-30, wherein: single Convolutional layers are designed on decomposed After the convolutional layer, or when the activation layer is placed in the single... After the convolutional layer.
32. The method according to any one of claims 1-31, wherein, A single NN-based SR is used to generate the output of the luminance component and the output of the chrominance component.
33. The method according to any one of claims 1-32, wherein, Two branches are designed in a single neural network, such that the first branch generates the output of the luminance component and the second branch generates the output of the chrominance component.
34. The method according to any one of claims 1-33, wherein, The inputs to the two branches come from different channels of the same feature map.
35. The method according to any one of claims 1-34, wherein, The two branches are designed with different network structures.
36. The method according to any one of claims 1-35, wherein, Each branch comprises multiple basic blocks, and: the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, modified linear units (ReLU), parameterized modified linear units (PReLU), or other layers used in the neural network, or the basic blocks use residual structures.
37. The method according to any one of claims 1-36, wherein, Compared to the branch that generates the luminance component, the branch that generates the chrominance component uses fewer convolutions, channels, or basic blocks.
38. The method according to any one of claims 1-37, wherein, An upsampling layer can be designed in either the luminance component branch or the chrominance component branch, or in both branches.
39. The method according to any one of claims 1-38, wherein, Pixel rearrangement is used as an upsampling layer.
40. The method according to any one of claims 1-39, wherein: The two branches share the same input, or the two branches are designed to have the same network structure, or each branch consists of multiple basic blocks, and the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, ReLU, PReLU, or other layers used in the neural network, or the basic blocks may or may not use residual structures, or a transposed convolution with a stride of 2 is used as an upsampling layer, or a convolutional layer is designed to follow the upsampling layer.
41. The method according to any one of claims 1-40, wherein, The reconstructed image and edge information are stitched together and then convolved.
42. The method according to any one of claims 1-41, wherein: The non-NN filter is a bilinear filter, a bicubic filter, or a Lanzos filter, or it is a concatenation of the upsampled chromaticity component of the predicted image with the luminance component of the predicted image.
43. The method according to any one of claims 1-42, wherein: The stripe type indicator is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered, or the stripe type value is a binary value that indicates whether the picture to be filtered includes intra-frame stripes.
44. The method according to any one of claims 1-43, wherein: The IPB information is first sliced or expanded into a two-dimensional array with the same size as the video unit to be filtered, or when the current CU block is in inter-frame prediction mode, the value of the IPB information is equal to A, where A is a constant value; or when the current CU block is in intra-frame prediction mode, the value of the IPB information is equal to B, where B is a constant value; or when the current CU block is in IBC prediction mode, the value of the IPB information is equal to C, where C is a constant value; or when the current CU block is a unidirectional prediction block and only L... When 0 is used for the current block, the value of the IPB information is equal to D, where D is a constant value; or when the current CU block is a one-way prediction block and only L1 is used for the current block, the value of the IPB information is equal to E, where E is a constant value; or when the current CU block is a one-way prediction block and both L0 and L1 are used for the current block, the value of the IPB information is equal to F, where F is a constant value; or when the current CU block is a two-way prediction block, the value of the IPB information is equal to G, where G is a constant value.
45. The method according to any one of claims 1-44, wherein: Convolutions share the same kernel size and different numbers of channels are assigned to each input, or the number of channels for side information inputs are different, or convolutions share the same number of convolution channels and different kernel sizes are assigned to each input, or The kernel size is used for convolution of part of the input, and The kernel size is used for convolution of the remaining input, and K represents an integer value greater than 1.
46. The method according to any one of claims 1-45, wherein: Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer is placed after each convolution, or Convolution is decomposed into Convolution and subsequent Combinations of convolutions, or Convolution is decomposed into Convolution and subsequent A combination of convolutions, and any activation layer is placed after each convolution, or Convolution is decomposed into Convolution and subsequent The convolutions are combined, and any activation layer is placed after each convolution, wherein... It is different from what is described positive integers, or Convolution is decomposed into Convolution and subsequent Combinations of convolutions, wherein, It is different from what is described positive integers, or Convolution is decomposed into Convolution and subsequent The convolutions are combined, and any activation layer is placed after each convolution, wherein... It is different from what is described Positive integers.
47. The method according to any one of claims 1-46, wherein: Any number Convolutional layers are designed on decomposed Before the convolutional layer, and the Convolutional layers are treated as multiple groups of two Combinations of convolutional layers, or as described Convolutional layers are simultaneously decomposed It is used before and after the convolutional layer.
48. The method according to any one of claims 1-47, wherein: In a neural network, a single branch is designed such that it generates both the luminance and chrominance components together, or the branch comprises multiple basic blocks, wherein the basic blocks are any combination of convolutional layers, fully connected layers, transformer layers, activation layers, ReLU, PReLU, or other layers used in the neural network, or the basic blocks use a residual structure, or an upsampling layer is designed for either the luminance component output or the chrominance component output, or an upsampling layer is designed for both the luminance component output and the chrominance component output, or pixel rearrangement is used as an upsampling layer, or a transposed convolution with a stride of 2 is used as the upsampling layer, or a convolutional layer is designed to follow the upsampling layer.
49. The method according to any one of claims 1-48, wherein: The output luminance component and output chrominance component generated by the single NN-based SR are used individually, or the output luminance component generated by the single NN-based SR is used and may not be used, or the output luminance component generated by the single NN-based SR may not be used and the output luminance component generated by the single NN-based SR is used.
50. The method according to any one of claims 1-49, wherein, The input image is enhanced, and the best encoding / decoding method is selected.
51. The method according to any one of claims 1-50, wherein: An enhancement is specified, or the enhancement is a rotation at any angle, or the rotation is applied at 90 degrees, or the rotation is applied at 180 degrees, or the rotation is applied at 280 degrees, or the enhancement is a flip, or the flip is applied vertically, or the flip is applied horizontally, or the enhancement is one or more combinations of rotation and flip at any angle.
52. The method according to any one of claims 1-51, wherein: The optimal encoding / decoding method is selected based on multi-pass compression, or the original input and all enhanced input images are compressed, and the optimal encoding / decoding method is selected based on rate-distortion optimization (RDO) selection of these compression results.
53. The method according to any one of claims 1-52, wherein: After the optimal encoding / decoding method is selected, the enhancement is transmitted via signaling, or the decoder generates an upsampled reconstruction based on the inverse operation of the enhancement transmitted via signaling.
54. The method according to any one of claims 1-53, wherein, The conversion includes encoding the visual media data into the bitstream.
55. The method according to any one of claims 1-53, wherein, The conversion includes decoding the visual media data from the bitstream.
56. An apparatus for processing video data, comprising: processor; and a non-transitory memory thereon having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-55.
57. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that when the computer-executable instructions are executed by a processor, the video codec apparatus performs the method according to any one of claims 1-55.
58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: determining the application side information as additional input to a neural network (NN)-based super-resolution (SR); and generating the bitstream based on the determination.
59. A method for storing a video bitstream, comprising: Determine the application of edge information as additional input to neural network (NN) based super-resolution (SR); The bit stream is generated based on the determination; And storing the bit stream in a non-transitory computer-readable recording medium.