Methods, apparatus and media for video processing
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2026-08-14
Smart Images

Figure CN122580685A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to neural network-based prediction. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between a current video block and a video bitstream, determining a first prediction for the current video block, the first prediction being determined based on a neural network-based intra-frame codec tool; determining a combined prediction of the first prediction and a second prediction for the current video block, the second prediction being determined based on an inter-frame codec tool; and performing a conversion based on the combined prediction. The method according to the first aspect of this disclosure enables the combination of predictions from a neural network-based intra-frame codec tool and predictions from an inter-frame codec tool.
[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a first prediction of a current video block, the first prediction being determined based on a neural network-based intra-frame encoding / decoding tool; determining a combined prediction of the first prediction and a second prediction of the current video block, the second prediction being determined based on an inter-frame encoding / decoding tool; and generating a bitstream based on the combined prediction.
[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a first prediction of a current video block, the first prediction being determined based on an intra-frame encoding / decoding tool based on a neural network; determining a combined prediction of the first prediction and a second prediction of the current video block, the second prediction being determined based on an inter-frame encoding / decoding tool; generating a bitstream based on the combined prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0011] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 An image of an 18×12 luminance CTU divided into 12 slices and 3 raster scan strips is shown (informative). Figure 5 An image showing an 18×12 luminance CTU divided into 24 segments and 9 rectangular stripes is presented (informative). Figure 6 The image shown is divided into 4 pieces, 11 bricks, and 4 rectangular strips (informative). Figures 7A to 7C An example of CTB across image boundaries is shown; Figure 8 An example of an encoder block diagram is shown; Figure 9 The preprocessing unit and the postprocessing unit are shown; Figure 10 The architecture of the CNN in filter set 0 is shown; Figure 11 The implementation of the CNN in filter set 0 is shown; Figure 12 Encoder optimization 2 is shown; Figures 13A to 13C The architecture of the CNN in filter set 1 is shown; Figure 14 A time-domain loop filter is shown; Figure 15A The parameter selection on the encoder side is shown; Figure 15B The parameter selection on the decoder side is shown; Figure 16 This demonstrates how intra-frame prediction patterns based on neural networks can be used to predict data from surrounding frames. Context of reference sample Predict the current situation piece ; Figure 17 It shows that it will revolve around the current piece Context of reference sample Decomposed into usable reference samples and unavailable reference points ; Figure 18 This shows the current [area] outlined in dashed lines. Intra-frame prediction mode signaling for Luminance CB; Figure 19 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 20 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0012] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0018] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. The codec images are codec representations of images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0023] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can, for example, select one of several codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0031] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0035] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0036] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0044] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring blocks (PBs) and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0050] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the block size of the frames(multiple) and / or stripes(multiple) used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially neighboring blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If needed, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while specific embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0056] 1. Brief Overview This disclosure relates to video codec technology. Specifically, it relates to neural network-based intra-frame prediction and inter-frame joint intra-frame prediction in image / video codec. It can be applied to existing video codec standards such as High Efficiency Video Codec (HEVC) and Multi-Functional Video Codec (VVC), or standards yet to be finalized (e.g., AVS3). It is also applicable to future video codec standards or video codecs.
[0057] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard (Johannes Ballé, Valero Laparra, and Eero P Simoncelli. 2016. End-to-end optimization of nonlinear transform codecs for perceptual quality. In PCS. IEEE, 1-5). Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform codecs. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploratory Model (JEM) (Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszár. 2017. Lossy image compression using a compression autoencoder. arXiv preprint arXiv:1703.00395 (2017)). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, aiming to reduce the bitrate by 50% compared to HEVC. VVC version 1 was finalized in July 2020.
[0058] The latest version of the VVC draft, namely the Multi-Functional Video Codec (Draft 10), can be found at the following location: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=10399 The latest reference software for VVC is called VTM, and it can be found at the following location: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags / VTM-10.0 The Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG is exploring potential neural network video codec technologies that go beyond the capabilities of VVC. This exploration is known as Neural Network-Based Video Coding (NNVC). Neural network-based (NN-based) codec tools are used to enhance or replace legacy modules in existing VVC designs. The implementation of NN-based tools in NNVC 4 is based on a small ad hoc deep learning (SADL) library.
[0059] The latest version of the draft description of the NNVC algorithm can be found at the following location: https: / / jvet-experts.org / doc_end_user / current_document.php?id=12563 NNVC-4.0 reference software is provided to demonstrate reference implementations of encoding and decoding processes, as well as training methods for neural network-based video encoding and decoding explored in JVET. The reference software can be accessed through the following methods: https: / / vcgit.hhi.fraunhofer.de / jvet-ahg-nnvc / VVCSoftware_VTM.
[0060] 2.1 Definition of Video Unit The image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular area of the image.
[0061] The sheet is divided into one or more bricks, each brick consisting of multiple CTU rows within the sheet.
[0062] A slice that is not divided into multiple bricks is also called a brick. However, a brick that is a proper subset of a slice is not called a slice.
[0063] A strip contains multiple slices of an image or multiple bricks of a slice.
[0064] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains multiple tiles that together form a rectangular area of the image. The tiles within the rectangular stripe are arranged in the order of the stripe's raster scan.
[0065] Figure 4 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.
[0066] Figure 5An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0067] Figure 6 An example of a picture partitioned into slices, tiles, and rectangular strips is shown, where the picture is partitioned into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the top - left slice contains 1 tile, the top - right slice contains 5 tiles, the bottom - left slice contains 2 tiles, and the bottom - right slice contains 3 tiles), and 4 rectangular strips.
[0068] 2.1.1. CTU / CTB Size In VVC, the CTU size signaled by the syntax element log2_ctu_size_minus2 in the SPS can be as small as 4×4.
[0069] 2.1.2. CTUs in a Picture Assume that the CTB / LCU size is indicated by M×N (usually M equals N, as defined in HEVC / VVC), and for a CTB located at the boundary of a picture (or a slice or a strip or other types of boundaries, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For those CTBs depicted as in Figures 7A to 7C , the CTB size is still equal to M×N. However, the lower boundary / right boundary of the CTB is outside the picture. In Figure 7A , K = M, L < N; in Figure 7B , K < M, L = N; and in Figure 7C , K < M, L < N.
[0070] 2.2. Coding and Decoding Streams of Typical Video Codecs Figure 8 An example of the encoder block diagram of VVC is shown, which includes three loop - filtering blocks: the de - block filter (DF), sample - adaptive offset (SAO), and ALF. Different from the DF that uses a predefined filter, SAO and ALF utilize the original samples of the current picture, reduce the mean - square error between the original samples and the reconstructed samples by adding an offset and by applying a finite - impulse response (FIR) filter respectively, where the coded - decoded side information signals the offset and the filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts caused by the previous stages.
[0071] 2.3. Neural - Network - Based Video Coding and Decoding (NNVC) 2.3.1. Neural - Network - Based Loop Filter Set 0 2.3.1.1. Pre - processing and Post - processing of Chrominance In filter set 0, filters with a single model are designed to process three components. Because luma and chroma have different resolutions, preprocessing and postprocessing steps are introduced to upsample and downsample the chroma component respectively, such as... Figure 9 As shown. During the resampling process, the nearest neighbor interpolation method is used.
[0072] 2.3.1.2. Neural Networks The network structure of a CNN filter is as follows: Figure 10 As shown. In addition to the reconstructed image (rec_yuv), additional side information is also fed into the network, such as the predicted image (pred_yuv), stripe QP, basic QP, and stripe type. In the residual block (ResBlock), the number of channels is first increased before the activation layer and then decreased after the activation layer. Specifically, K and M are set to 64 and 160, respectively, and the number of Resblocks is set to 32.
[0073] 2.3.1.3. Combination with traditional filters Figure 11 This illustrates the implementation of a CNN in filter set 0. For example... Figure 11 As shown, the reconstructed samples prior to DBK are fed into a CNN-based filter (CNNLF), and then the final filtered samples are generated by mixing the results of CNNLF and SAO. This mixing process can be briefly described as follows:
[0074] For the blending weights, there are four candidates: 1, 0.75, 0.5, and adaptive weights. The derivation of the adaptive weights is based on the least squares method. If adaptive weights are chosen, the blending weights are transmitted via signal transmission for each color component in the strip header.
[0075] 2.3.1.4. Mode Selection CNN filters can be turned on / off at both the CTU and strip levels. For each enable type, there are four hybrid modes. Therefore, there are nine modes to be evaluated by the RDO at the encoder. The finally selected mode will be transmitted via signal in the strip header.
[0076] Table 1. Parameter selection for filter set 0
[0077] 2.3.1.5. Basic QP Adjustment The basic QP is fed into the CNN filter, such as Figure 10As shown. To improve adaptability, an offset can be added to the base QP at the strip level (the adjusted base QP is used as the input to the NN filter). The offset candidates are {-5, 5}. For example, given an offset of -5, for the current strip, the actual input base QP of the filter becomes (BaseQP-5).
[0078] Encoder method The proposed encoder filters only one of every four CTUs during the process of selecting the optimal basic QP offset, thus saving encoding time. Figure 12 Encoder optimization 2 is shown. (As shown) Figure 12 As shown, only the shaded CTU is considered to calculate the distortion using different BaseQP candidates {BaseQP, BaseQP-5, BaseQP+5}. After selecting the candidate with the minimum cost, the encoder applies the optimal offset to the base QP to adjust the distortion of the remaining CTUs. Figure 12 The non-shadowed CTU in the image is filtered.
[0079] 2.3.1.6. Encoder-only optimization To more accurately estimate the rate-distortion (RD) cost of utilizing the integrated NN-based loop filter, the encoder-only NN filter is involved in the segmentation decision process. In the segmentation mode decision, the distortion between the NN-filtered samples and the original samples is calculated, and then the optimal segmentation mode is selected based on the calculated distortion to make the segmentation decision more accurate. To reduce complexity, only a small number of ResBlocks are used in the network architecture (see Section 3.1.2). The NN filter in the RDO process is implemented using SADL with integer (int) 16 precision. This encoder-only NN tool is disabled by default.
[0080] 2.3.1.7. Reasoning Details SADL (see Section 1.3) is used to perform inference for CNN filters. It supports both floating-point and fixed-point implementations. In the fixed-point implementation, static quantization is used, and weights and feature maps are represented with int16 precision. Table 2 provides network information for the inference phase.
[0081] Table 2. Network information for filter set 0 during the inference phase
[0082] 2.3.2. Loop Filter Set Based on Neural Networks 1 2.3.2.1. Neural Network for Luminance Component Filter set 1 contains two conventional networks, one for luminance and one for chrominance.
[0083] The input to the luminance network includes reconstructed luminance samples (rec), predicted luminance samples (pred), boundary intensities (bs), QP, and block type (IPB). The number of feature maps and residual blocks are set to 96 and 8, respectively. Figures 13A to 13C The architecture of the CNN in filter set 1 is shown. The structure of the brightness network is as follows. Figures 13A to 13C As shown. Figure 13A The head of the luminance network is shown, where the inputs are combined to form the inputs for the next part of the network. y . Figure 13B The output of the head is shown for the k-th residual block (k=0...7). y Feeded to input z 0 =y In the first residual block, and output z 1 is then fed into another similar residual block. Figure 13C This shows that the output of the last residual block is fed into this final part of the network.
[0084] 2.3.2.2. Neural Networks for Chromaticity Components Luminance information is used as additional input to the loop filter for chroma. Considering that luminance has a higher resolution than chroma in the YUV 4:2:0 format, features are first extracted from luminance and chroma separately. The luminance features are then downsampled and concatenated with the chroma features. The input to the chroma network includes reconstructed luminance samples (recY), reconstructed chroma samples (recUV), predicted chroma samples (predUV), boundary intensity (bsUV), and QP. Regarding the network backbone, the chroma component uses the same backbone as the luminance component.
[0085] 2.3.2.3. Time-domain filter Filter set 1 contains additional loop filters, i.e., time-domain filters, which obtain co-occurrence blocks from the first image in the two reference image lists to improve performance. Figure 14 A time-domain loop filter is shown. (Example) Figure 14 As shown, two co-located blocks are directly concatenated and fed into the network. When the temporal filtering feature is enabled, the temporal filter is applied to the luminance component of the image in the three highest temporal layers, while the regular luminance and chrominance filters are used for other purposes. By default, this temporal filtering feature is disabled.
[0086] Figure 14 The time-domain loop filter is shown. Only the header is shown; the other parts are shown separately. Figures 13B-13C The same in {column 0, column 1} refers to the cosine samples from the first image in the two reference image lists.
[0087] 2.3.2.4. Adaptive Inference Granularity The granularity of filter determination and parameter selection depends on the resolution and QP. Given a higher resolution and a larger QP, determination and selection will be performed over a larger area.
[0088] 2.3.2.5. Parameter Selection Each strip or block can determine whether a CNN-based filter is applied. Once a CNN-based filter is determined to be applied to a strip / block, it can be further determined which conditional parameter to choose from a candidate list including three candidates derived from the QP. Let the sequence-level QP be denoted as q, and the candidate list include the conditional parameters {Param_1, Param_2, Param_3}. For lower temporal layers, Param_1 = q, Param_2 = q 5. Param_3 = q 10. For higher time-domain layers, Param_1 = q, Param_2 = q 5. Param_3 = q 5. In other words, the third candidate is different in different time domain layers.
[0089] The selection process is based on the rate-distortion cost on the encoder side. If necessary, the on / off control indication and condition parameter index are transmitted via signaling in the bitstream. Figure 15A and Figure 15B This diagram illustrates parameter selection on both the encoder and decoder sides. All blocks in the current frame are first processed using three conditional parameters. Then, five costs, Cost_0, ..., Cost_5, are calculated and compared to achieve optimal rate-distortion performance. In Cost_0, CNN-based filters are disabled for all blocks. In Cost_i, {i = 1, 2, 3}, the parameter Param_i is applied to all blocks. In Cost_4, different blocks can prefer different parameters, and for each block, information about whether to use a CNN-based filter or which parameters to use is transmitted via signal transmission. On the decoder side, whether to use a CNN-based filter or which parameters to use for a block is based on... Figure 15B The Param_Id of the bitstream parsing is shown.
[0090] Note that for the intra-frame configuration, parameter selection is disabled, while filter on / off control is still retained. Shared conditional parameters are used for both chroma components to alleviate the worst-case burden on the decoder side. Additionally, the maximum number of conditional parameter candidates can be specified on the encoder side.
[0091] 2.3.2.6. Residual Scaling When a neural network (NN) filter is applied to reconstruct an image, a scaling factor is derived and transmitted via signal transmission for each color component in the strip header. The derivation is based on the least squares method. The difference between the input samples and the NN-filtered samples (residuals) is scaled by the scaling factor before being added to the input samples.
[0092] 2.3.2.7. Combination with deblocking filters To enable the combination with deblocking, the input samples used in residual scaling are the output of the deblocking filter. The residual scaling process is shown below, where... and These refer to the outputs of NN filtering and deblocking filtering, respectively.
[0093] =
[0094] 2.3.2.8. Encoder Optimization Only Unlike NNVC-2.0, EncDbOpt also has AI configuration enabled.
[0095] To better estimate the rate-distortion (RD) cost when using neural network (NN) filters, the proposed encoder introduces NN-based filtering into the rate-distortion optimization (RDO) process for segmentation mode selection. Specifically, the refinement distortion is calculated by comparing the NN-filtered samples with the original samples. The segmentation mode with the minimum rate-refinement distortion cost is selected as the optimal segmentation mode. To reduce complexity, several fast algorithms are applied. First, the NN model is simplified by using a smaller number of residual blocks. Second, parameter selection is not allowed for the NN filtering in the RDO process. Third, the proposed technique is only applied to encoder / decoder units with a height and width no greater than 64. The NN filters used in the RDO process are also implemented via SADL using fixed-point-based computation. This NN-based encoder-only method is disabled by default.
[0096] 2.3.2.9. Reasoning Details SADL (see Section 1.3) is used to perform inference for CNN filters. It supports both floating-point and fixed-point implementations. In the fixed-point implementation, static quantization is used, and weights and feature maps are represented with int16 precision. Table 3 provides network information for the inference phase.
[0097] Table 3. Network information of filter set 1 during the inference phase
[0098] 2.3.3. Intra-frame prediction based on neural networks 2.3.3.1. Neural Network Inference The neural network-based intra-frame prediction mode comprises seven neural networks, each predicting... Blocks of different sizes. Predicted size: The block neural network is represented as ,in Collect its parameters. For a given piece , Using the block located above OK A reference sample point and its left side List The context composed of reference samples Preprocessed version to provide Applying post-processing produce Prediction ,See Figure 16 . Figure 16 This demonstrates how intra-frame prediction patterns based on neural networks can be used to predict data from surrounding frames. Context of reference sample Predict the current situation piece .exist Figure 16 In the example, and .also, Return two indexes and . This represents the index that characterizes the LFNST kernel index, and when When applying DCT-2 horizontally and vertically to the residuals predicted by the neural network, are the master transform coefficients transposed? ,See Figure 16 .also, Provide an index of the VVC intra-prediction modes (planar intra-prediction mode, DC intra-prediction mode, or directional intra-prediction mode). Among them, from surrounding Reference sample The prediction best represents ,See Figure 16 .
[0099] if :
[0100] otherwise: if :
[0101] otherwise:
[0102] if :
[0103] otherwise:
[0104] if , .otherwise, .
[0105] if , .otherwise, .
[0106] 2.3.3.2. Preprocessing and Postprocessing 2.3.3.2.1. Preprocessing of the current block's context Figure 16 The “preprocessing” shown consists of the following four steps.
[0107] • from Subtract Available reference samples mean ,See Figure 17 . Figure 17 It shows that it will revolve around the current piece Context of reference sample Decomposed into usable reference samples and unavailable reference points In the case shown, the number of unavailable reference samples reaches its maximum value.
[0108] • If the neural network predicting the current block is floating-point, then the context... Multiply the reference sample points in , It refers to the internal bit depth, i.e., in VVC. Otherwise, the context Multiply the reference sample points in , This indicates the input quantizer.
[0109] • All unavailable reference points ,See Figure 17 They were all set to .
[0110] • Flatten the context generated in the previous step to produce a size of [size missing]. vector .
[0111] 2.3.3.2.2. Post-processing of neural network predictions Figure 16 The "post-processing" described in the text includes processing data of a size of [missing information]. vector Remodeling to a height of and width are The rectangle, the result of the reshaping is divided by Add the mean of available reference samples in the context of the current block. And limited to Therefore, post-processing can be summarized as follows: .
[0112] 2.3.3.3. Adaptive Derivation of the MPM List When creating a list of MPMs for a given luminance CB, if the "left" luminance CB is predicted via a neural network-based intra-frame prediction mode, the neural network-based mode index can be obtained from the data returned during the prediction of the "left" luminance CB. The index is replaced and becomes a candidate index to be added to the MPM list. Similarly, if the "above" luminance CB is predicted via a neural network-based intra-frame prediction mode, the neural network-based mode index can be obtained from the index returned during the prediction of the "above" luminance CB. Replace it and make it a candidate index to be inserted into the MPM list.
[0113] 2.3.3.4. Signaling for Intra-Frame Prediction Mode Based on Neural Networks 2.3.3.4.1. Signaling for Intra-Frame Prediction Mode Based on Neural Networks in Luminosity The position of the current top-left pixel in the current luminance channel is: of In the luminance CB, intra-predictive mode signaling in luminance is divided into two cases.
[0114] • if ,but nnFlag It appears in the intra-prediction mode signaling in the brightness. nnFlag This means that the neural network-based intra-frame prediction mode is selected to predict the current brightness CB and END. nnFlag This means that the intra-frame prediction mode based on the neural network was not selected to predict the current luminance CB, so the luminance is represented as The standard intra-frame prediction mode signaling is applicable, see Figure 18 . Figure 18 This shows the current [area] outlined in dashed lines. Intra-prediction mode signaling for the luminance CB. The coordinates of the top-left pixel of this CB are... The binary value of nnFlag is shown in coarse gray. Here, , , ,and .
[0115] Otherwise, regular intra-frame prediction mode signaling in brightness. Applicable.
[0116] Note: In " In the case where the context of the current luminance channel CB exceeds the boundary of the current luminance channel, i.e. Then, the intra-frame prediction based on the neural network is replaced by the plane.
[0117]
[0118] .
[0119] 2.3.3.4.2. Signaling for Intra-Frame Prediction Mode Based on Neural Networks in Chroma The position of the current top-left pixel in the current chroma channel is: of In chroma CB, intra-frame prediction mode signaling in chroma is divided into two cases.
[0120] • If the luminance CB, which is in the same position as the chrominance CB, is predicted by an intra-frame prediction mode based on a neural network: o If Then DM becomes an intra-frame prediction mode based on neural networks.
[0121] Otherwise, DM is set to a plane.
[0122] • Otherwise: o If ,but nnFlagChroma It appears in the intra-predictive mode signaling in chroma. nnFlagChroma It is placed before the DM flag in the decision tree of the intra-predictive mode signaling in chroma. nnFlagChroma This means that the neural network-based intra-frame prediction mode is selected to predict the current pair of chroma CB and END. nnFlagChroma This means that if the neural network-based intra-prediction mode is not selected to predict the current pair of chroma CBs, then the regular intra-prediction mode signaling in the chroma is restored from the DM flag.
[0123] Otherwise, the standard intra-frame prediction mode signaling in chroma applies.
[0124] Note: In " In the case of and in " In the case of "", if the context of the current chroma CB exceeds the boundary of the current chroma channel, i.e. Then, the intra-frame prediction based on the neural network is replaced by the plane.
[0125] 2.3.3.5. Transformation of Context and Neural Network Prediction For a given block, if Therefore, the intra-frame prediction mode based on neural networks must predict this block, but the intra-frame prediction mode based on neural networks does not include... This is possible. In this case, the context of the current block can be... Figure 16 The vertical downsampling factor is referred to in the "preprocessing" step. and / or level downsampling factor And / or transpose. Then, the prediction of the current block can be... Figure 16 The step referred to as "post-processing" is followed by transposition and / or vertical upsampling factor. and / or level upsampling factor The context of the current block and the predicted transpose, and The selected neural networks, belonging to the intra-frame prediction mode based on neural networks, are used for prediction, as shown in Table 4.
[0126]
[0127] Table 4: Current situation to be predicted The context of the block and the decision of the predicted transpose of the block, The value and of Values, and belonging to each The neural network-based intra-frame prediction pattern prediction 2.3.4. Small Ad hoc deep learning (SADL) library SADL (Small Ad hoc Deep Learning Library) is a small, header-only library for neural network inference. SADL provides both floating-point and integer-based inference capabilities. Neural network inference in NNVC is based on SADL.
[0128] The following table summarizes the framework features.
[0129] Table 5. Characteristics of SADL
[0130] The NNVC repository uses SADL as a submodule, pointing to this repository: https: / / vcgit.hhi.fraunhofer.de / jvet-ahg-nnvc / sadl.
[0131] The documentation is available in the doc directory of the repository.
[0132] 3. Problem Current NN-based intra-frame prediction suffers from the following problems: 1. NN-based intra-frame prediction is used as an additional prediction mode in intra-frame prediction. However, NN-based intra-frame prediction is not fully considered in inter-frame prediction, such as joint intra-inter-frame prediction (CIIP).
[0133] 4. Detailed Solution The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.
[0134] In the following discussion, a video unit can be a sequence, picture, strip, slice, brick, sub-picture, CTU / CTB, CTU / CTB line, one or more CU / CB, one or more CTU / CTB, one or more VPDU (Virtual Pipeline Data Unit), or a sub-region within a picture / strip / slice / brick. A parent video unit represents a unit larger than a video unit. Typically, a parent unit will contain several video units. For example, when the video unit is a CTU, the parent unit can be a strip, a CTU line, multiple CTUs, etc.
[0135] 1. NN-based intra-frame prediction (NN-intra) can be combined with any inter-frame prediction method.
[0136] a. In one example, inter-frame prediction can be intra-inter-frame joint prediction (CIIP), merge, skip, direct, affine, geometric segmentation (GEO), history-based motion vector prediction (HMVP), sub-block temporal motion vector prediction (SbTMVP), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), etc.
[0137] 2. NN-intra-frame and inter-frame prediction can be combined into new prediction modes.
[0138] a. In one example, the new prediction pattern is one of the existing prediction patterns, such as CIIP.
[0139] 3. NN-intra-frame and inter-frame predictions can be combined by weighting the predicted samples. The output of the prediction is the weighted samples.
[0140] a. In one example, the weighting factors for NN-intra-frame and inter-frame predictions are (1,3) or (3,1) or (1,1) or (2:2).
[0141] b. In one example, the weighting factors for NN-based intra-frame and inter-frame predictions can be generated by an NN-based network.
[0142] c. In one example, the weighting factor for NN-intra-frame and inter-frame prediction is the output signal of the intra-frame prediction network based on NN.
[0143] d. In one example, the weighting factor depends on the statistics of the codec.
[0144] i. In one example, it depends on the pattern type of the neighboring blocks.
[0145] e. In one example, the weighting factor is the same as that of the CIIP pattern.
[0146] 4. NN-intra-frame and inter-frame prediction can be combined for all components, or only the luminance component, or only the chrominance component.
[0147] a. In one example, when NN-intraframes are used for combining the luminance components, the regular intraframe mode is used for combining with interframe predictions for the chrominance components.
[0148] i. In one example, the regular intra-frame mode is either planar mode or direct mode.
[0149] b. In one example, when NN-intraframes are used for the combination of luminance components, NN-intraframes are used for the combination with inter-frame predictions for chrominance components.
[0150] c. In one example, if NN-intraframe is not available, other intraframe modes are used in combination with interframe prediction in CIIP.
[0151] i. In one example, when the direct mode is used for the chroma component and the NN-frame is used for the luma component, it can also be used when the NN-frame is not available for the chroma component.
[0152] ii. In one example, it can also be used when the luminance component is unavailable within an NN-frame.
[0153] iii. In one example, it can also be used when the chroma component is not available within an NN-frame.
[0154] iv. In one example, when the above situation is encountered, the planar pattern is used for composition.
[0155] d. In one example, the above items are applicable to CIIP mode.
[0156] 5. Examples 5.1. Example 1 The use of NN-intra-frame prediction in CIIP mode is proposed. In CIIP mode, NN-intra-frame prediction is used only for the luma component, and the plane is used for the chroma component. The weighting factors for NN-intra-frame prediction and inter-frame prediction remain the same as in CIIP mode.
[0157] Figure 19 A flowchart of a method 1900 for video processing according to an embodiment of the present disclosure is shown. Method 1900 is implemented during the conversion between video units of a video and a bitstream of a video.
[0158] At box 1910, for the conversion between the current video block and the video bitstream, a first prediction for the current video block is determined. The first prediction is determined based on an intra-frame encoding / decoding tool based on a neural network (NN).
[0159] At box 1920, a combined prediction of the first and second predictions for the current video block is determined. The second prediction is determined based on inter-frame encoding / decoding tools. That is, inter-frame prediction and NN-based intra-frame prediction can be combined. The combined prediction can be referred to as a combination of inter-frame prediction and NN-based intra-frame prediction.
[0160] At box 1930, the transformation is performed based on the combined predictions.
[0161] Method 1900 enables the combination of predictions from neural network-based intra-frame codec tools with predictions from inter-frame codec tools. In this way, codec efficiency and codec effectiveness can be improved.
[0162] In some embodiments, the inter-frame coding / decoding tool includes at least one of the following: inter-frame and intra-frame joint prediction (CIIP) coding / decoding tool, merge coding / decoding tool, skip coding / decoding tool, direct coding / decoding tool, affine coding / decoding tool, geometric segmentation coding / decoding tool, history-based motion vector prediction (HMVP) coding / decoding tool, sub-block temporal motion vector prediction (SbTMVP) coding / decoding tool, bidirectional optical flow (BDOF) coding / decoding tool, or decoder-side motion vector refinement (DMVR) coding / decoding tool.
[0163] In some embodiments, determining the combined prediction includes an encoding / decoding mode. For example, the encoding / decoding mode may be an inter-frame / intra-frame joint prediction (CIIP) mode. That is, this combination of inter-frame prediction and NN-based intra-frame prediction can be considered as a new prediction mode, or as a CIIP prediction mode.
[0164] In some embodiments, the combined predictions are determined by weighting the prediction samples of the current video block. For example, the prediction samples are included in at least one of a first prediction or a second prediction.
[0165] In some embodiments, the weighting factor for the predicted samples includes one of the following: (1, 3), (3, 1), (1, 1), or (2:2). For example, a weighting factor (1, 3) means that the weighting factor for the first prediction is 1 and the weighting factor for the second prediction is 3.
[0166] In some embodiments, the weighting factors for the predicted samples are determined by a neural network-based network.
[0167] In some embodiments, the weighting factor for the predicted samples is determined based on a neural network-based network of a neural network-based intra-frame codec tool.
[0168] In some embodiments, the weighting factor for the predicted samples depends on the statistics associated with the transformation.
[0169] In some embodiments, the statistics include the mode types of neighboring blocks of the current video block.
[0170] In some embodiments, the weighting factor for the predicted samples is the same as the weighting factor for the inter-frame and intra-frame joint prediction (CIIP) codec tool.
[0171] In some embodiments, the combined prediction is determined based on at least one of the luminance or chrominance components of the first and second predictions.
[0172] In some embodiments, if a neural network-based intra-frame encoding / decoding tool is used for the combined prediction of the luminance component, then the intra-frame mode is used to determine the combined prediction of the chrominance component.
[0173] In some embodiments, the intra-frame mode includes at least one of a planar mode or a direct mode.
[0174] In some embodiments, if a neural network-based intra-frame codec tool is used for the combined prediction of the luminance component, then the neural network-based intra-frame codec tool is used for the combined prediction of the chrominance component.
[0175] In some embodiments, if a neural network-based intra-codec tool is not available to determine the combined prediction, at least one intra-mode different from the neural network-based intra-codec tool is used to determine the intra-prediction for the current video block. Intra-prediction can be used in inter-frame intra-prediction joint prediction (CIIP).
[0176] In some embodiments, if a direct mode is used for the combined prediction of the chroma components and a neural network-based intra-frame encoding / decoding tool is used for the combined prediction of the luma components, then at least one intra-frame mode can be used to determine the intra-frame prediction.
[0177] In some embodiments, if a neural network-based intra-frame encoding / decoding tool is not available to determine the combined prediction of the luminance components, then at least one intra-frame mode is used to determine the intra-frame prediction.
[0178] In some embodiments, if a neural network-based intra-frame codec tool is not available to determine the combined prediction of chroma components, then at least one intra-frame mode is used to determine the intra-frame prediction.
[0179] In some embodiments, at least one intra-frame mode includes a planar mode.
[0180] In some embodiments, the combined predictions are applied to the inter-frame and intra-frame joint prediction (CIIP) mode.
[0181] In some embodiments, the conversion includes encoding the current video block into a bitstream.
[0182] In some embodiments, the conversion includes decoding the current video block from the bitstream.
[0183] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, a first prediction of a current video block is determined. The first prediction is determined based on an intra-frame encoding / decoding tool based on a neural network. A combined prediction of the first prediction and a second prediction of the current video block is determined. A second prediction is determined based on an inter-frame encoding / decoding tool. The bitstream is generated based on the combined predictions.
[0184] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, a first prediction of a current video block is determined. The first prediction is determined based on an intra-frame encoding / decoding tool based on a neural network. A combined prediction of the first prediction and a second prediction of the current video block is determined. The second prediction is determined based on an inter-frame encoding / decoding tool. A bitstream is generated based on the combined predictions. The bitstream is stored in a non-transitory computer-readable recording medium.
[0185] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0186] Item 1. A method for video processing, comprising: for a conversion between a current video block and a bitstream of the video, determining a first prediction of the current video block, the first prediction being determined based on an intra-frame encoding / decoding tool based on a neural network; determining a combined prediction of the first prediction and a second prediction of the current video block, the second prediction being determined based on an inter-frame encoding / decoding tool; and performing a conversion based on the combined prediction.
[0187] Item 2. The method according to Item 1, wherein the inter-frame coding / decoding tool includes at least one of the following: inter-frame intra-frame joint prediction (CIIP) coding / decoding tool, merge coding / decoding tool, skip coding / decoding tool, direct coding / decoding tool, affine coding / decoding tool, geometric segmentation coding / decoding tool, history-based motion vector prediction (HMVP) coding / decoding tool, sub-block temporal motion vector prediction (SbTMVP) coding / decoding tool, bidirectional optical flow (BDOF) coding / decoding tool, or decoder-side motion vector refinement (DMVR) coding / decoding tool.
[0188] Item 3. The method according to Item 1 or 2, wherein the determination of the combined prediction includes a coding / decoding mode, the coding / decoding mode including an inter-frame / intra-frame joint prediction (CIIP) mode.
[0189] Item 4. The method according to any one of items 1-3, wherein the combined prediction is determined by weighting the prediction samples of the current video block, the prediction samples being included in at least one of the first prediction or the second prediction.
[0190] Item 5. According to the method described in Item 4, the weighting factor of the predicted sample includes one of the following: (1,3), (3,1), (1,1), or (2:2).
[0191] Item 6. The method according to Item 4, wherein the weighting factor of the predicted sample is determined by a neural network-based network.
[0192] Item 7. The method according to Item 4, wherein the weighting factor of the predicted sample is determined based on the neural network of the neural network-based intra-frame codec tool.
[0193] Item 8. The method according to Item 4, wherein the weighting factor of the predicted sample depends on the statistics associated with the transformation.
[0194] Item 9. The method according to Item 8, wherein the statistical information includes the mode type of the neighboring blocks of the current video block.
[0195] Item 10. The method according to Item 4, wherein the weighting factor of the predicted sample is the same as the weighting factor of the inter-frame and intra-frame joint prediction (CIIP) codec tool.
[0196] Item 11. The method according to any one of Items 1 to 10, wherein the combined prediction is determined based on at least one of the luminance component or chromaticity component of the first prediction and the second prediction.
[0197] Item 12. The method according to Item 11, wherein in response to the neural network-based intra-frame encoding / decoding tool being used for the combined prediction of the luminance component, an intra-frame mode is used to determine the combined prediction of the chrominance component.
[0198] Item 13. The method according to Item 12, wherein the intra-frame mode includes at least one of a planar mode or a direct mode.
[0199] Item 14. The method according to Item 12, wherein the neural network-based intra-frame codec tool is used for the combined prediction of the luminance component, and the neural network-based intra-frame codec tool is used for the combined prediction of the chrominance component.
[0200] Item 15. The method according to Item 12, wherein if it is determined that the neural network-based intra-codec tool is not available for determining the combined prediction, at least one intra-mode different from the neural network-based intra-codec tool is used to determine the intra-prediction of the current video block, the intra-prediction being used in inter-frame intra-prediction joint prediction (CIIP).
[0201] Item 16. The method according to Item 15, wherein in response to a direct mode being used for the combined prediction of the chroma components, and the neural network-based intra-frame encoding / decoding tool being used for the combined prediction of the luma components, the at least one intra-frame mode being used to determine the intra-frame prediction.
[0202] Item 17. The method according to Item 15, wherein if it is determined that the neural network-based intra-frame codec tool is not available for determining the combined prediction of the luminance component, the at least one intra-frame mode is used to determine the intra-frame prediction.
[0203] Item 18. The method according to Item 15, wherein the at least one intra-frame mode is used to determine the intra-frame prediction based on the determination that the neural network-based intra-frame codec tool is not available for determining the combined prediction of the chroma components.
[0204] Item 19. The method according to any one of items 16-18, wherein the at least one intra-frame mode includes a planar mode.
[0205] Item 20. The method according to any one of items 1-19, wherein the combined prediction is applied to an inter-frame and intra-frame joint prediction (CIIP) mode.
[0206] Item 21. The method according to any one of items 1-20, wherein the conversion includes encoding the current video block into the bitstream.
[0207] Item 22. The method according to any one of items 1-20, wherein the conversion includes decoding the current video block from the bitstream.
[0208] Item 23. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-22.
[0209] Item 24. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-22.
[0210] Item 25. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining a first prediction of a current video block of the video, the first prediction being determined based on an intra-frame encoding / decoding tool based on a neural network; determining a combined prediction of the first prediction and a second prediction of the current video block, the second prediction being determined based on an inter-frame encoding / decoding tool; and generating the bitstream based on the combined prediction.
[0211] Item 26. A method for storing a bitstream of video, comprising: determining a first prediction of a current video block of the video, the first prediction being determined based on an intra-frame codec tool based on a neural network; determining a combined prediction of the first prediction and a second prediction of the current video block, the second prediction being determined based on an inter-frame codec tool; generating the bitstream based on the combined prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0212] Example device Figure 20 A block diagram of a computing device 2000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0213] It should be understood that, Figure 20 The computing device 2000 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0214] like Figure 20 As shown, computing device 2000 includes general-purpose computing device 2000. Computing device 2000 may include at least one or more processors or processing units 2010, memory 2020, storage unit 2030, one or more communication units 2040, one or more input devices 2050, and one or more output devices 2060.
[0215] In some embodiments, the computing device 2000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2000 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0216] Processing unit 2010 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2000. Processing unit 2010 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0217] Computing device 2000 typically includes various computer storage media. Such media can be any media accessible by computing device 2000, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2030 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2000.
[0218] The computing device 2000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 20 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0219] The communication unit 2040 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 2000 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, the computing device 2000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0220] Input device 2050 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2060 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2040, computing device 2000 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2000 can also communicate with one or more devices that enable a user to interact with computing device 2000, or any device that enables computing device 2000 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via input / output (I / O) interface (not shown).
[0221] In some embodiments, some or all components of the computing device 2000 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and work together to achieve the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (e.g., the Internet) using suitable protocols. For example, a cloud computing provider offers applications via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers, or directly installed or otherwise installed on client devices.
[0222] In embodiments of this disclosure, computing device 2000 can be used to implement video encoding / decoding. Memory 2020 may include one or more video codec modules 2025 having one or more program instructions. These modules can be accessed and executed by processing unit 2010 to perform the functions of the various embodiments described herein.
[0223] In an example embodiment performing video encoding, input device 2050 may receive video data as input 2070 to be encoded. The video data may be processed, for example, by video codec module 2025 to generate an encoded bitstream. The encoded bitstream may be provided as output 2080 via output device 2060.
[0224] In an example embodiment performing video decoding, input device 2050 may receive an encoded bitstream as input 2070. The encoded bitstream may be processed, for example, by video codec module 2025 to generate decoded video data. The decoded video data may be provided as output 2080 via output device 2060.
[0225] While this disclosure has been specifically shown and described with reference to exemplary embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. Such changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, a first prediction of the current video block is determined, the first prediction being determined based on a neural network-based intra-frame encoding / decoding tool; Determine a combined prediction of the first and second predictions for the current video block, wherein the second prediction is determined based on inter-frame encoding / decoding tools; and The transformation is performed based on the combined predictions.
2. The method according to claim 1, wherein the inter-frame encoding / decoding tool comprises at least one of the following: Inter-frame and intra-frame joint prediction (CIIP) encoding and decoding tools, Merge encoding / decoding tool Skip the codec tools, Direct encoding / decoding tools Affine encoding / decoding tools Geometric segmentation encoding and decoding tools History-based motion vector prediction (HMVP) encoding and decoding tools, Sub-block temporal motion vector prediction (SbTMVP) encoding and decoding tool, Two-way optical flow (BDOF) codec tools, or Decoder-side Motion Vector Refinement (DMVR) encoding / decoding tool.
3. The method according to claim 1 or 2, wherein the determination of the combined prediction includes a coding / decoding mode, the coding / decoding mode including an inter-frame / intra-frame joint prediction (CIIP) mode.
4. The method according to any one of claims 1 to 3, wherein the combined prediction is determined by weighting the prediction samples of the current video block, the prediction samples being included in at least one of the first prediction or the second prediction.
5. The method according to claim 4, wherein the weighting factor of the predicted sample includes one of the following: (1, 3), (3, 1), (1, 1), or (2: 2).
6. The method of claim 4, wherein the weighting factor of the predicted sample is determined by a neural network-based network.
7. The method of claim 4, wherein the weighting factor of the predicted sample is determined based on the neural network of the neural network-based intra-frame codec tool.
8. The method of claim 4, wherein the weighting factor of the predicted sample depends on statistical information associated with the transformation.
9. The method of claim 8, wherein the statistical information includes the mode types of neighboring blocks of the current video block.
10. The method of claim 4, wherein the weighting factor of the predicted sample is the same as the weighting factor of the inter-frame and intra-frame joint prediction (CIIP) codec tool.
11. The method according to any one of claims 1 to 10, wherein the combined prediction is determined based on at least one of the luminance component or chromaticity component of the first prediction and the second prediction.
12. The method of claim 11, wherein in response to the neural network-based intra-frame encoding / decoding tool being used for the combined prediction of the luminance component, the intra-frame mode is used to determine the combined prediction of the chrominance component.
13. The method of claim 12, wherein the intra-frame mode includes at least one of a planar mode or a direct mode.
14. The method of claim 12, wherein the neural network-based intra-frame codec tool is used for the combined prediction of the luminance component, and the neural network-based intra-frame codec tool is used for the combined prediction of the chrominance component.
15. The method of claim 12, wherein if it is determined that the neural network-based intra-codec tool is not available for determining the combined prediction, at least one intra-mode different from the neural network-based intra-codec tool is used to determine the intra-prediction of the current video block, the intra-prediction being used in inter-frame intra-prediction (CIIP).
16. The method of claim 15, wherein in response to the direct mode being used for the combined prediction of the chroma components, and the neural network-based intra-frame codec tool being used for the combined prediction of the luma components, the at least one intra-frame mode being used to determine the intra-frame prediction.
17. The method of claim 15, wherein if it is determined that the neural network-based intra-frame codec tool is not available for determining the combined prediction of the luminance component, the at least one intra-frame mode is used to determine the intra-frame prediction.
18. The method of claim 15, wherein if it is determined that the neural network-based intra-frame codec tool is not available for determining the combined prediction of the chroma components, the at least one intra-frame mode is used to determine the intra-frame prediction.
19. The method according to any one of claims 16 to 18, wherein the at least one intra-frame mode includes a planar mode.
20. The method according to any one of claims 1 to 19, wherein the combined prediction is applied to an inter-frame and intra-frame joint prediction (CIIP) mode.
21. The method according to any one of claims 1 to 20, wherein the conversion comprises encoding the current video block into the bitstream.
22. The method of any one of claims 1 to 20, wherein the conversion comprises decoding the current video block from the bitstream.
23. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 22.
24. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1 to 22.
25. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: A first prediction is determined for the current video block of the video, the first prediction being determined based on a neural network-based intra-frame encoding / decoding tool; Determine a combined prediction of the first and second predictions for the current video block, wherein the second prediction is determined based on inter-frame encoding / decoding tools; and The bitstream is generated based on the combined predictions.
26. A method for storing a bitstream of video, comprising: A first prediction is determined for the current video block of the video, the first prediction being determined based on a neural network-based intra-frame encoding / decoding tool; Determine a combined prediction of the first and second predictions for the current video block, wherein the second prediction is determined based on inter-frame encoding / decoding tools; The bitstream is generated based on the combined predictions; as well as The bitstream is stored in a non-transitory computer-readable recording medium.