Intercoding Using Deep Learning in Video Compression

By integrating luma-chroma networks, attention layers, and improved training methods, the inter-frame coding efficiency and complexity of deep learning-based video coding are enhanced, achieving significant gains in bitrate performance and computational efficiency.

JP2025522769APending Publication Date: 2025-07-17DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024576398
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-11
Filing Date
2023-06-23
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing deep learning-based video coding technologies, such as DLVC, face challenges in achieving optimal inter-coding efficiency and complexity, particularly in handling the correlation between luma and chroma components in YUV420 format, and in effectively utilizing temporal and spatial correlations in motion and residual coding.

Method used

Implementing luma-chroma integrated motion and residual coding networks, incorporating attention layers, temporal motion prediction, cross-domain fusion, and weighted motion-compensated inter-prediction, along with improved training methods like large motion training and temporal distance modulation loss, to enhance inter-frame coding efficiency.

Benefits of technology

Improves coding efficiency by 10-14% in Y channel and 0-26% in U or V channels, reduces computational complexity, and enhances bitrate performance by 1.5-2.5% through optimized neural network training and adaptive block usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522769000001_ABST
    Figure 2025522769000001_ABST
Patent Text Reader

Abstract

A method, system, and bitstream syntax for inter-frame coding using an end-to-end neural network used in image and video compression are described. The inter-frame coding method includes luma-chroma integrated motion compensation of YUV pictures, luma-chroma integrated residual coding of YUV pictures, use of an attention layer, activation of a temporal motion prediction network for motion vector prediction, use of an intersection area network that combines motion vectors and residual information for motion vector decoding, use of an intersection area network for decoding residuals, use of weighted motion-compensated inter-prediction, and includes one or more of use of only temporal features, only spatial features, or both temporal and spatial features in entropy decoding. A method for improving training of a neural network for inter-frame coding is also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to images. More specifically, embodiments of the present invention relate to inter-coding using deep learning in video compression.

Background Art

[0002] In 2020, the MPEG group within the International Standardization Organization (ISO), in cooperation with the International Telecommunications Union (ITU), published the first version of the Versatile Video Coding standard (VVC), also known as H.266. More recently, the same joint group (JVET) and experts in still image compression (JPEG) have begun working on the development of a next-generation coding standard that provides coding performance superior to existing image and video coding technologies. As part of this research, coding technologies based on artificial intelligence and deep learning are also being considered. As used herein, the term "deep learning" refers to a neural network having at least three layers, preferably more than three layers.

[0003] As understood by the inventors herein, improved techniques for coding images and videos based on neural networks are described herein.

[0004] The approaches described in this section are approaches that could have been pursued, but are not necessarily approaches that have been previously devised or pursued. Therefore, unless otherwise indicated, none of the approaches described in this section should be assumed to be eligible as prior art solely by virtue of being included in this section. Similarly, the problems identified with respect to one or more of the approaches should not be assumed to be recognized in any of the prior art based on this section, unless otherwise indicated.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 5A

Figure 5B

Figure 6

Figure 7

[0006] Embodiments of the present invention are represented, by way of example and not limitation, in the figures of the accompanying drawings. In the figures, the same reference numerals refer to similar elements.

[0007] Exemplary embodiments regarding inter-frame coding in the case of using a neural network in the coding of images and videos are described herein. In the following description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments of the present invention. It will be apparent, however, that the various embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in detail so as not to unnecessarily obscure, obfuscate, or make the embodiments of the present invention difficult to understand.

[0008] [Overview] The exemplary embodiments described herein relate to the coding of images and videos using neural networks. In an embodiment, a processor receives a coded video sequence and a high-level syntax indicating that inter-coding adaptation is enabled for decoding a current picture, and the processor parses the high-level syntax to extract inter-coding adaptation parameters, decodes the current picture based on the inter-coding adaptation parameters to generate an output picture, and the inter-coding adaptation parameters include a luma-chroma integration motion compensation flag indicating that a joint luma-chroma motion compensation network is used in decoding when an input picture is within a YUV color region, a luma-chroma integration residue coding flag indicating that a joint luma-chroma residue network is used in decoding when an input picture is within a YUV color region, an attention layer flag indicating that an attention network layer is used in decoding, a temporal motion prediction flag indicating that a temporal motion prediction network is used for motion vector prediction in decoding, a cross-region motion vector flag indicating that a cross-region network that combines a motion vector and residual information is used for decoding the motion vector in decoding, a cross-region residue flag indicating that a cross-region network that combines a motion vector and residual information is used for decoding the residual in decoding, a spatio-temporal entropy flag indicating whether entropy decoding uses only spatial features, only temporal features, or a combination of spatial and temporal features, and includes one or more of the above.

[0009] In a second embodiment, in a system including a processor that trains a neural network for inter-frame coding, the processor Large motion training performed using random P-frame skipping from 1 to n - 1 in the case of a training sequence with a total of n pictures, and A time-distance modulation loss that calculates a rate-distortion loss such as Loss = w × lambda × MSE + Rate One or more of these can be used, Rate represents the achieved bitrate, MSE measures the distortion between the original picture and the corresponding reconstructed picture, and the weight parameter w

Number

[0010] In a third embodiment, a method for processing an uncompressed video frame with one or more neural networks is provided. The method Generating motion vector information and spatial map information (α) of an uncompressed input video frame (x t ) based on the continuity of uncompressed input video frames including the uncompressed input video frame, and Generating a motion-compensated frame based at least on the reference frame and the motion compensation network used to generate the motion vector information

Number

[0011] [Exemplary Coding Model Using Deep Learning] Deep learning-based approaches for image and video compression are becoming increasingly common and are an active research area. Figure 1 depicts an example of a basic, deep learning-based network (reference [1]). It includes some of the basic components (e.g., motion compensation, motion estimation, residual coding, etc.) found in conventional codecs such as advanced video coding (AVC), high-efficiency video coding (HEVC), versatile video coding (VVC), etc. The main difference is that all of these components use neural network (NN)-based approaches such as a motion vector (MV) decoder network, a motion compensation (MC) network, a residual decoder network, etc. The framework also includes some encoder-only components such as an optical flow network, an MV encoder network, a residual encoder network, quantization, etc. Such a framework is commonly referred to as an end-to-end deep-learning video coding (DLVC) framework.

[0012] Note that this end-to-end deep learning (DL) network does not have an inverse quantization block (inverse Q), unlike conventional encoder architectures. Such an end-to-end network does not require an inverse Q. This is because quantization based on simple half-rounding of latents is performed at the encoder size, eliminating the need for any inverse Q on the decoder side. The network is trained for different lambdas (different QPs) to generate one model per lambda (e.g., loss = lambda × MSE + Rate).

[0013] Compared with conventional coding schemes, the latest DLVC approach has similar coding performance for images, but there is still a significant gap regarding inter-coding when compared with VVC. The embodiments described herein focus on improving the training, coding efficiency, and coding complexity of neural networks regarding inter-frame (or inter) coding.

[0014] [YUV4:2:0 Coding] In a typical DLVC implementation, the framework of FIG. 1 operates on images within the RGB domain. Considering the correlation between chroma components, it may be more efficient to operate in a luma-chroma space such as YUV, YCbCr, etc. within the 4:2:0 region (simply denoted as YUV420 without limitation). Here, 4:2:0 indicates that the chroma components are subsampled by a factor of 2 in both horizontal and vertical resolutions compared to luma.

[0015] To operate in the YUV420 domain, several changes have been proposed to enable more efficient YUV420 coding. Since the motion of luma and chroma is highly correlated, in an embodiment, luma and chroma motion estimation and motion coding are jointly performed by using a modified YUV optical flow network and an MV encoder Net-MV decoder Net respectively. However, the motion compensation and residual coding of the luma component and chroma component of YUV420 can be processed in a number of ways as follows: · Use separate motion compensation (MC) networks for luma and chroma, or use a luma-chroma integrated MC network; · Use separate residual coding networks for luma and chroma, or use a luma-chroma integrated residual coding network.

[0016] [Separate Motion Compensation (MC) Network and Integrated Motion Compensation (MC) Network for YUV420 Coding] MC networks designed for RGB images assume that all image channels have the same dimension. In the case of a YUV420 inter-frame, as shown in Figure 2A, a separate MC network can be devised to match the sizes of the Y channel and the UV channel. However, this has additional complexity, the coupling information present in the Y channel and the UV channel is not effectively utilized, and there is also a risk that artifacts may occur in the reconstructed image due to slightly different motion corrections of the channels. The input to the luma MC network within the separate MC network of Figure 2A is

Number

Number

Number

Number

[0017] For example, the luma-chroma integrated MC network shown in FIG. 2B can effectively utilize the interdependence on the condition that Y and UV references as well as the warped frame channels are properly handled. The input to the luma-chroma integrated MC network of FIG. 2B is

Number

Number

[0018] Figure 2C represents an exemplary embodiment of a neural network for luma-chroma integrated MC. A typical MC neural network (see References [1] and [7]) includes a first convolutional layer including residual blocks acting on the spatial dimension of the current frame, followed by a) a series of average pooling layers that reduce the spatial dimension of the predicted frame features by half, and b) residual blocks. The predicted frame features of lower spatial dimension are then processed using a series of residual blocks, upsampled to enhance the quality of inter prediction, and added directly to higher dimensional features. In the case of YUV420, since the chroma components have half the resolution of the luma, the motion compensation of the luma and chroma inter prediction components is performed integrally by merging the chroma channels with appropriate pooling layers of the luma where their resolutions match. The proposed method realizes reduction of computational cost and improvement of performance, and at the same time reduces the memory usage.

[0019] In the luma-chroma integrated MC network of Figure 2C, the linearly interpolated frames of luma and chroma are independently and partially processed using convolutional and residual blocks in the first stage (before 205 and 210). Then, due to the chroma having half the resolution of the luma, the chroma inter prediction features (205) are added to the luma inter prediction after the first luma pooling layer (215). This ensures that the luma prediction and chroma prediction are then processed jointly, reducing complexity compared to separate MC networks and also leveraging the inter-channel dependencies. The chroma inter prediction features are separated from the combined inter prediction features before the final upsampling layer (235) and processed separately from (240) to output the final motion-compensated chroma inter prediction.

Number

[0020] As used in Figure 2C, the term Conv(K,C,S) represents a convolutional network with a K×K kernel, C output channels, and a stride S (where S = 1 means no upsampling or downsampling). In the notation, the number of outputs from a given stage is assumed to be equal to the number of inputs to the next stage, so the number of inputs is not explicitly shown. For example, in column 240, Conv(3,64,1) is followed by Conv(3,2,1). That is, the last layer, Conv(3,2,1), receives 64 input channels from the previous layer, Conv(3,64,1), and outputs two channels corresponding to the chroma MC prediction output [Number] outputs two channels corresponding to it.

[0021] [Separate Residual Coding Network and Integrated Residual Coding Network for YUV420 Coding] Similar to the consideration of the MC network, the luma and chroma residuals of the inter-frame can be coded separately or jointly. Separate residual coding can improve the coding performance of chroma. However, separate residual coding may increase complexity and may increase coding overhead if the cross-correlation between the luma residual channel and the chroma residual channel is not effectively utilized. The luma / chroma separate residual coding network is new for inter-frame coding. The luma-chroma integrated residual coding network can effectively utilize the interdependence of the residuals while simultaneously reducing the complexity of the residual network and the entropy coding. The current integrated residual coding architecture is based on References [5] and [6].

[0022] [Addition of Attention Layers in Intercoding] Figure 3A represents an example of a process pipeline (300) for video coding using a four-layer neural network architecture for encoding and decoding latent features (Reference [7]). As used herein, "latent features" (also referred to as "latent characteristics") or "latent variables" (also referred to as "latent variables") are features or variables that are not directly observable but are inferred from other observable features or variables, for example, by processing directly observable variables. In the coding of images and videos, a "latent space" (also referred to as a "latent space") may refer to a compressed representation of data in which similar data points are closer together. In video coding, examples of latent features include transform coefficients, residuals, motion representations, syntax elements, model information., and other such representations. In the context of neural networks, the latent space is useful for learning data features and finding a simpler representation of image data for analysis.

[0023] As shown in FIG. 3A, considering the input image x(302) at the input h×w resolution, in the encoder (300E), the input image is processed by a series of convolutional neural network blocks (also called convolutional networks or convolutional blocks) (305, 310, 315, 320) each followed by a non-linear activation function. In each such layer (which may include multiple sub-layers of convolutional networks and activation functions), its output is typically reduced (e.g., usually reduced by a factor of 2 or more, which is usually called "stride", stride = 1 means no downsampling, stride = 2 means downsampling by a factor of 2 in each direction, etc.). For example, using stride = 2, the output of the L1 convolutional network (305) becomes h / 2×w / 2. The last layer (e.g., 320) generates the output latent coefficient y(322), which is further quantized (Q) and entropy encoded (e.g., by an arithmetic encoder AE) before being sent to the decoder (300D). Hyper-prior networks and spatial context model networks (not shown) are also used to generate a probabilistic model of the latent (y).

[0024] In the decoder (300D), the process is reversed. After arithmetic decoding (AD),

Number

Number

[0025] As shown in Figure 3B, in an embodiment, it has been proposed to add "attention blocks" for the encoding and decoding of P-frames and B-frames. Attention blocks (e.g., block 355) are used to emphasize certain data more than other data. As an example, an attention block can be added after two layers. In other embodiments, attention blocks can be added after each layer, but the improvement in performance may not justify the increase in complexity.

[0026] The reason for using adaptive blocks is that conventional video codecs gain significant benefits from their block-level adaptation to local image / video characteristics. Thus, DLVC should also benefit from local adaptation. An attention block is one way to locally adapt a layer by weighting the filter response using spatially varying weights that are learned end-to-end with the filter. Attention blocks can be applied to, for example, an MV network, and / or a residual network, and / or an MC network. Their use in a specified neural network can be signaled to the decoder using high-level syntax elements. Examples of such syntax elements are described later in this specification. Using the proposed architecture, experimental results with YUV420 data have shown that the BD-rate improves by 10 - 14% in Y and 0 - 26% in U or V.

[0027] [Temporal Motion (Flow) Prediction] In the current P-frame model and B-frame model, the total bits spent on coding motion information and residual information account for most of the total bitrate. The motion field is correlated both temporally and spatially. In Reference [1], the motion field generated by the optical flow network utilizes spatial correlation but not temporal correlation. In Reference [4], multiple previous decoded frames are used as inputs to explore temporal information. In the embodiments, it is proposed to explore temporal correlation with DLVC.

[0028] In an exemplary embodiment, it is proposed to use temporal information based on a flow prediction network that takes as input the motion fields of one or more previous frames. Experimental results show that by using two frames, a good trade-off (a 2% BD-rate gain in translation) between complexity and performance can be achieved. Figure 4A shows an example of an NN for temporal MV prediction.

[0029] As shown in Figure 4A, the proposed NN (400) includes a flow buffer, a convolutional 2D network, a series of ResBlock64 layers (405), and a final convolutional 2D network, which are used for motion prediction of the current frame using the decoded flow of the previously decoded frames.

[0030] Figures 4A and 4C show examples of applying flow prediction with temporal, delta, and motion vector coding for the cases of P-frames and B-frames respectively. Figure 4B shows the decoded flow of the past three (L0) reference frames

Number

Number

Number

Number

Number

[0031] Even if the architecture of FIG. 4A provides a coding gain, the prediction may be suboptimal. This is because when a significant amount of motion exists, the previous two motion fields and the current frame may not spatially correspond to each other, and in a network with limited field sizes, it may be difficult to infer the spatial correspondence and internally align them to properly predict the current motion field. To address this limitation, in other embodiments, it has been proposed to align the motion fields before providing them as input to the prediction network. When the motion fields at the previous two instants and the current instant are represented as

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

[0032] [Cross-domain Fusion for Motion and Residual Coding] In Reference [1], motion and residual coding are performed independently. In an embodiment, it is proposed to utilize the potential cross-correlation between motion features and residual features. For example, the discontinuity of motion at object boundaries can be used to more effectively code residual features.

[0033] In an embodiment, it is proposed to use cross-domain fusion for motion vector (MV) coding. In the embodiment shown in FIG. 5A, the reconstructed sample (502) of the previous frame is used to enhance MV coding (505) in the encoder using, for example, optical flow (such as block 400). In the decoder, the previous frame residual latency value (508) is further used for MV compensation (MC) of the current frame to utilize the motion interdependence with respect to the residual. The motion vector decoder block (510) applies cross-domain fusion using the residual latency (508) and motion vector latency (506) of the previous frame. The motion vector encoder (505) applies cross-domain fusion based on the previous frame image (502) as an additional input to the motion vector encoding process. The fusion is performed in the latency domain in the decoder and in the spatial domain in the encoder. This fusion method attempts to utilize any interdependence of the motion of the current frame with respect to the strength of the current image or residual image. As an example, non-zero residuals at object boundaries are likely to coincide with motion boundaries, which can help improve the coding efficiency of motion information.

[0034] In other embodiments, cross-region fusion may be applied with residual coding. As shown in FIG. 5B, the reconstructed motion vectors can be used to derive the residual coding (520). In the decoder, the residual decoder (525) utilizes both the motion vector latency (507) and the residual latency. The residual decoder block (525) applies cross-region fusion by using the motion vector latency (507) of the current frame. The residual encoder (520) applies cross-region fusion using the reconstructed motion of the current frame as an additional input to the residual encoding process. The fusion is performed in the latency region in the decoder and in the spatial region in the encoder. This fusion method attempts to utilize any interdependence of the residual of the current frame with respect to the motion of the current frame within the same region. As an example, a change in the motion field at an object boundary is likely to coincide with a non-zero residual, which can help improve the coding efficiency of the residual information.

[0035] [Entropy Coding Based on Temporal and Spatial Priors] In embodiments, it is desirable to enable an entropy NN model to use features from previous frames or from the spatial neighborhood. In reference [8], the core idea is that the entropy model estimates the spatio-temporal redundancy in the latent space rather than at the pixel level, thereby significantly reducing the complexity of the framework.

[0036] In inter-frame coding, the residual intensity map undergoes approximately the same motion as the current image. Since the encoder CNN network is shift-invariant, even if the size is reduced by the downsampling ratio performed in the network layer, the latent feature map is also transformed with approximately the same motion. When warping the latent map of the residual of the previous frame with an appropriately downsampled and scaled image motion field, the current latent to be transmitted can be appropriately predicted. The entropy model of the latent can be conditioned on the predicted latent in addition to the hyperprior latent and the latent of the already decoded current frame. This should result in a significant reduction in the number of bits required to transmit the residual latent. Figure 6 represents an example of the proposed entropy model with the addition of a temporal entropy model. The spatial context model uses the decoded neighboring latent features of the current frame to estimate the spatial model parameter φ t to estimate the spatial model parameter φ [Number] and the hyperprior decoded features [Number] are used to estimate the hyperprior parameter (Ψ t ), and the latent features of the previously decoded frame [Number] are warped using the decoded motion of the current frame used to estimate the temporal prior feature γ t (e.g., by using linear interpolation). These three features are jointly used to estimate the entropy model parameters of a Gaussian or Laplace or multiple mixture model, such as the mean and variance of the latent of the next frame of the current frame. One point to note is the motion field of the current frame [Number] should be scaled and downsampled to match the latency's spatial resolution of [Number]

[0037] [Improvement of Training] To improve the efficiency of inter-frame coding, the following training procedure is proposed: 1) Large motion training: To handle large motions appropriately, it is necessary to add large motion training using random P-frame skipping from 1 to n - 1, where n represents the total number of frames in the training clip (e.g., n = 7).

[0038] 2) Temporal distance modulation loss: The idea is to use higher weights for more distant P-frames: In the case of temporal distance modulation loss, the rate-distortion (RD) can be formulated as follows: Loss = w × lambda × MSE + Rate Here, Rate represents the achieved bitrate (bits per pixel), and MSE measures the L2 loss between the original picture and the corresponding reconstructed picture. The weight parameter "w" is initialized based on the temporal distance (t) of the inter-frame using the following distance-based weighting: [Number] Assuming that i represents the number of iterations (e.g., from 1 to 200,000), this weight w i is the same for each frame within a group of pictures (GOP) and monotonically increases from 0 to 1 over a cycle of 200,000 iterations.

[0039] 3) MV Entropy Modulation Loss: The idea is to give higher weights to low-probability latencies (samples that are difficult to encode) compared to high-probability latencies. It is motivated by the focal loss in object detection (Reference [3]). In object detection, there is always an imbalance between background samples and foreground samples, and the network often confuses background and foreground. In Reference [3], the authors proposed a solution to this problem. They applied certain weights to the cross-entropy loss to weight more hard samples that are background samples in the image.

[0040] In an embodiment, the formulation of the weighted entropy loss is as follows: Entropy Loss = -log(p t ) Modulation Entropy Loss = (1 - p t ) b × (-log(p t )) Here, pt is the estimated probability of the latency symbol t , and b monotonically decreases from 5.0 to 0.0 over a cycle of 200,000 iterations. When b = 0, the modulation entropy loss decreases to the normal entropy loss.

[0041] The overall coding gain due to the improved training procedure is from about 1.5% to 2.5%.

[0042] [Syntax Example] The proposed tool can be sent from the encoder to the decoder using high-level syntax (HLS) that can be part of a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), picture header (PH), slice header (SH), or as part of supplementary metadata such as supplementary enhancement information (SEI). An example syntax is shown in Table 1. Alternatively, such signaling is not required if a specific architecture or tool is pre-determined and known to both the encoder and the decoder. [Table 1]

[0043] An inter_coding_adaptation_enabled_flag equal to 1 specifies that inter-coding adaptation is enabled for the decoded picture. An inter_coding_adaptation_enabled_flag equal to 0 specifies that inter-coding adaptation is not enabled for the decoded picture.

[0044] A joint_LC_MC_NN_enabled_flag equal to 1 specifies that the luma-chroma integrated MC network is used to decode the signal within the YUV region. A joint_LC_MC_NN_enabled_flag equal to 0 specifies that a separate MC network is used to decode the signal within the YUV region.

[0045] A joint_LC_residue_NN_enabled_flag equal to 1 specifies that the luma-chroma integrated residue network is used to decode signals within the YUV region. A joint_LC_residue_NN_enabled_flag equal to 0 specifies that separate residue networks are used to decode signals within the YUV region.

[0046] An attention_layer_enabled_flag equal to 1 specifies that the attention layer is enabled for the decoded picture. An attention_layer_enabled_flag equal to 0 specifies that the attention layer is not enabled for the decoded picture.

[0047] An attention_layer_MV_enabled_flag equal to 1 specifies that the attention layer is enabled for MV decoding. An attention_layer_MV_enabled_flag equal to 0 specifies that the attention layer is not enabled for MV decoding.

[0048] An attention_layer_residue_enabled_flag equal to 1 specifies that the attention layer is enabled for residue decoding. An attention_layer_residue_enabled_flag equal to 0 specifies that the attention layer is not enabled for residue decoding.

[0049] A temporal_motion_prediction_idc equal to 0 specifies that the temporal motion prediction net module is not used to decode the motion vector. A temporal_motion_prediction_idc equal to 1 specifies that the temporal motion prediction net module including simple concatenation is used to decode the motion vector. A temporal_motion_prediction_idc equal to 2 specifies that the temporal motion prediction net module including warping the reference picture is used to decode the motion vector.

[0050] The number obtained by adding 1 to num_ref_pics_minus1 specifies the number of reference pictures used in the temporal motion prediction net module.

[0051] A cross_domain_mv_enabled_flag equal to 1 specifies that the cross-domain net is enabled to decode the motion vector. A cross_domain_mv_enabled_flag equal to 0 specifies that the cross-domain net is not enabled to decode the motion vector.

[0052] A cross_domain_residue_enabled_flag equal to 1 specifies that the cross-domain net is enabled to decode the residue. A cross_domain_residue_enabled_flag equal to 0 specifies that the cross-domain net is not enabled to decode the residue.

[0053] A temporal_spatio_entropy_idc equal to 0 specifies that neither the temporal feature nor the spatial feature is used for entropy decoding within an inter-frame. A temporal_spatio_entropy_idc equal to 1 specifies that only the spatial feature is used for entropy decoding within an inter-frame. A temporal_spatio_entropy_idc equal to 2 specifies that only the temporal feature is used for entropy decoding within an inter-frame. A temporal_spatio_entropy_idc equal to 3 specifies that both the temporal feature and the spatial feature are used for entropy decoding within an inter-frame.

[0054] [Weighted Motion-Compensated Inter-Prediction] FIG. 7 represents an example of a process for weighted motion-compensated prediction according to an embodiment. Compared with FIG. 1, FIG. 7 represents the following changes: replacing the MV encoder network with an MV+Inter-Weight Map encoder network, replacing the MV decoder with an MV+Inter-Weight Map decoder network, and adding a “Blend Inter” network.

[0055] The motivation for this architecture is to allow both intra-coding and inter-coding in an inter-frame without explicit signaling of the binary intra / inter flag used in conventional block-based coding. Explicit binary mode signaling was easy in conventional block-based codecs, but may not be easy or effective for deep learning-based codecs that code features that partially overlap in the latency region.

[0056] Instead of using the binary intra / inter flag, the basic idea here is to code an explicit spatial weight map along with the motion information. These weights are, as shown in FIG. 7, for residual coding (r t) followed by motion-compensated (MC) inter-prediction samples

Number

Table 2

[0057] The process of weighted motion-compensated prediction can be expressed as follows:

Number

[0058] Note that the motion latency carries information for both the compressed flow and the spatial weight map.

[0059] The motion compression network output also includes the spatial weight map (α or α t ) used to mix motion compensation before residual compression. The residual is the difference between the original frame and the motion-compensated pixels scaled by α at the pixel resolution level.

[0060] When α is 1 (unity), since the decoded flow and the reliability of the corresponding MC inter prediction are high, the residual is purely inter-coded. When α is 0 (zero), since the decoded flow and the reliability of the corresponding MC inter prediction are low, the residual is purely intra-coded. The mixing of intra information and inter information can be coded when α is between 0 and 1, based on the quality of the inter prediction. This is similar to the combined intra-inter prediction (CIIP) used in conventional codecs such as VVC.

[0061] The network is trained in an end-to-end manner for a common rate-distortion (RD) loss function, e.g., loss = lambda × MSE + Rate, during which network parameters for optimal coding of motion information, alpha weights, and residual information are learned using a stochastic gradient descent algorithm such as the ADAM optimizer. This network is trained for large video datasets such as Vimeo-90k using a batch size of 4, 8, or 16. A network trained on a generalized video dataset may not fully understand the selection of motion information, alpha weights, and residual information that is optimal for RD for the actual source content during testing. In embodiments, this can be mitigated by content-specific encoder optimization, e.g., by overfitting the encoder network or the coded latent for a given source video using an iterative refinement procedure. This can help optimize the alpha weights, motion information, and residual information to minimize the RD loss for a given content under test, even as the complexity of the encoder increases.

[0062] To accurately train and code the spatial weight map, additional inputs are included (extended) in the motion and alpha compression network (e.g., MV + inter-weight map encoder Net), such as the use of a previously reconstructed reference frame, the use of the current original input frame, and the use of a warped reference frame generated using uncompressed optical flow motion vectors, etc. The warping process is not shown in the figure for simplicity.

[0063] The residual compression architecture remains the same as in the previous version of DLVC (see, e.g., Figure 1). For the mixing of motion compensation, as discussed earlier, other variations can also be done to improve training by passing this information to the decoder instead of the warped frame, for example, by using the argument of the syntax variable mv_aug_type. For example, the following options are available: 1. mv_aug_type = prev_com_res: The previous frame (reference frame) and the residual latency of the reference frame are used as extended inputs 2. mv_aug_type = input: The source frame (input), the reference frame, and the residual latency of the reference frame are used as extended inputs 3. mv_aug_type = warp: The source frame (input), the reference frame, the warped reference frame (using uncompressed flow), and the residual latency of the reference frame are used as extended inputs 4. mv_aug_type = mc: The source frame (input), the reference frame, the motion-compensated reference frame (using uncompressed flow), and the residual latency of the reference frame are used as extended inputs

[0064] Experimental results show that the proposed scheme has the potential to improve the compression efficiency by at least 2% and up to 5% compared to DLVC v.4.4 (Reference [1]), depending on the class of the test images.

[0065] [References] Each of the reference documents cited in this specification is hereby incorporated by reference in its entirety. [1]Guo Lu et al., "DVC: An end-to-end deep video compression framework", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2019. [2]Z. Guo et al., “Soft then Hard: Rethinking the Quantization in Neural Image Compression”, Proceedings of ICML 2021, arXiv:2104.05168v1, 12 April, 2021. [3]Tsung-Yi Lin et al., “Focal Loss for Dense Object Detection”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, arXiv:1708.02002v2, 7 Feb., 2018. [4]Z. Hu et al., “FVC: A New Framework towards Deep Video Compression in Feature Space”, IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2021, arXiv: 2105.09600v1, 20 May, 2021. [5]A. K. Singh et al., “A Combined Deep Learning based End-to-End Video Coding Architecture for YUV Color Space”, CVPR 2021, arXiv:2104.00807v1, April 1, 2021. [6]H. E. Egilmez et al., “Transform Network Architectures for Deep Learning based End-to-End Image / Video Coding in Subsampled Color Spaces”, arXiv:2103.01760v1, 27 Feb., 2021. [7]A. Mohananchettiar et al., “Multi-level latent fusion in neural networks for image and video coding”, US Provisional Patent Application No. 63 / 257388, filed on Oct. 19, 2021 [8]Z. Sun et al., "Spatiotemporal entropy model is all you need for learned video compression", arXiv:2104.06083 (2021). [9]Z. Cheng et al., “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules”, CVPR 2020, arXiv: 2001.01568v3, 30 March, 2020.

[0066] [Example Computer System Implementation] Embodiments of the present invention may be implemented by a computer system, an electronic circuit, and a system, microcontroller, field-programmable gate array (FPGA), or other configurable or programmable logic device (PLD) such as an integrated circuit (IC) device, a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions regarding inter-frame coding using a neural network for coding of images and videos, such as those described herein. The computer and / or IC may calculate any of the various parameters or values related to inter-frame coding using a neural network for coding of images and videos, as described herein. Embodiments of images and videos may be implemented in hardware, software, firmware, and various combinations thereof.

[0067] One implementation of the present invention has a computer processor that executes software instructions that cause the processor to perform the method of the present invention. For example, one or more processors included in, for example, a display, an encoder, a set-top box, a transcoder, etc. may implement a method for inter-frame coding using a neural network for coding of images and videos as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory and tangible medium carrying a set of computer-readable signals that, when executed by a data processor, include instructions that cause the data processor to perform the method of the present invention. The program product according to the present invention may take any of a variety of non-transitory and tangible forms. The program product may include, for example, magnetic data storage media including floppy (registered trademark) disks and hard disk drives, optical data storage media including CD ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0068] Components (e.g., software modules, processors, assemblies, devices, circuits, etc.) have been mentioned above, but unless otherwise indicated, any reference to such a component (including reference to a "means") shall be construed to include any component that performs the function of the recited component (e.g., is functionally equivalent), including components that are not structurally equivalent to the disclosed structure that performs the function in the exemplary embodiments of the present invention described.

[0069] [Equivalents, Extensions, Alternatives, and Others] Examples of embodiments related to inter - frame coding using neural networks for the coding of images and videos have been described. In the above - mentioned specification, embodiments of the present invention have been described with reference to a number of specific details that may vary from implementation to implementation. Accordingly, the only and exclusive indicator of what the invention is and what the applicant intends to be the invention is a series of claims issued from this application, including the specific form in which those claims are issued and subsequent amendments. The definitions explicitly set forth herein for the terms included in the claims shall govern the meaning of the terms used in the claims. Accordingly, limitations, elements, characteristics, features, advantages, or attributes not explicitly recited in the claims shall not limit the scope of those claims in any way. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a limiting sense.

[0070] [Cross - reference to related applications] This application claims the benefit of priority to Indian Patent Provisional Application No. 202241037461, filed on June 29, 2022, and Indian Patent Provisional Application No. 202341026932, filed on April 11, 2023, each of the prior Indian Patent Provisional Applications being incorporated by reference in its entirety.

Claims

1. A method for processing a coded video sequence with one or more neural networks, comprising: receiving a high-level syntax indicating that inter-coding adaptation is enabled for decoding of the current picture; parsing the high-level syntax to extract inter-coding adaptation parameters; decoding the current picture based on the inter-coding adaptation parameters to generate an output picture ; wherein the inter-coding adaptation parameters include a luma-chroma integrated motion compensation enable flag indicating that a luma-chroma integrated motion compensation network is used in decoding when the input picture is within the YUV color region; a luma-chroma integrated residual coding enable flag indicating that a luma-chroma integrated residual network is used in decoding when the input picture is within the YUV color region; an attention layer enable flag indicating that an attention network layer is used in decoding; a temporal motion prediction enable flag indicating that a temporal motion prediction network is used for motion vector prediction in decoding; a cross-region motion vector enable flag indicating that a cross-region network for combining motion vectors and residual information is used for decoding motion vectors in decoding; a cross-region residual enable flag indicating that a cross-region network for combining motion vectors and residual information is used for decoding residuals in decoding; a spatio-temporal entropy flag indicating whether entropy decoding uses only spatial features, only temporal features, or a combination of spatial and temporal features ; and one or more of the above. A method.

2. When luma-chroma integrated motion compensation is enabled, the motion-compensated luma component 【Number 1】 and chroma component 【Number 2】 of the current picture are generated using a luma-chroma integrated motion compensation network having an input comprising 【Number 3】 the decoded motion of the current picture, 【Number 4】 the luma component of the previous reference picture, 【Number 5】 the bilinearly interpolated luma prediction picture represented thereby,the chroma component of the previous reference picture,and the bilinearly interpolated chroma prediction component represented thereby using the downsampled and downscaled chroma motion; 【Number 6】 The method according to claim 1. 【Number 7】

3. 【Number 8】 ​ ​ ​ ​ When an attention network layer is used, when decoding a P picture or a B picture, an attention block layer is inserted between two transposed convolutional layers, and each transposed convolutional layer includes a transposed convolutional layer with upsampling and a subsequent non-linear activation block. The method according to claim 1.

4. The attention block layer is inserted after two consecutive transposed convolutional layers, and there is no attention block between them, or is inserted after each transposed convolutional layer. The method according to claim 3.

5. When a temporal motion prediction network is used for motion vector prediction, the flow prediction neural network includes a flow buffer that receives the decoded flow information, a subsequent first convolutional 2D network, a subsequent sequence of one or more ResBlock64 layers, and a subsequent second convolutional 2D network. The method according to claim 1.

6. In the case of a P picture, generating the output motion of the current picture in decoding 【Number 9】 is using the decoded flow motions of three past reference pictures as input to the flow prediction neural network to generate a first output, 【Number 10】 receiving the quantized delta motion value from the encoder, processing the quantized delta motion value using a motion vector decoder network to generate an input motion prediction value, and adding the first output to the input motion prediction value to generate the output motion vector of the current picture. And having generating the quantized delta motion value in the encoder To generate a second output, the current picture (X t ) and a reference picture 【Number 11】 by using a motion estimation block, subtracting the first output from the second output to generate a delta output, and processing the delta output by a motion vector coding network and subsequent quantization to generate the delta motion value. And having The method according to claim 5.

7. In the case of a B picture, generating the output motion of the current picture in decoding 【Number 12】 is using the decoded flow motions of two past reference pictures and one future reference picture as input to the flow prediction neural network to generate a first output, 【Number 13】 receiving the quantized delta motion value from the encoder, ​ processing the quantized delta motion values using a motion vector decoder network to generate an input motion prediction value; adding the first output to the input motion prediction value to generate an output motion vector of the current picture; and; generating the quantized delta motion values in the encoder includes: To generate the second output, the current picture (X t ) and the reference picture 【Number 14】 using a motion estimation block; subtracting the first output from the second output to generate a delta output; processing the delta output by a motion vector coding network followed by quantization to generate the delta motion values. and; The method according to claim 5. **Claim 8** There is a warping network before the input to the flow prediction neural network, and the warping network includes: 【Number 15】 a first network having a motion vector input for generating an output to be inverse warped by; 【Number 16】 a concatenation network that concatenates with the output of the first network to generate a concatenation network output; 【Number 17】 a second network that generates the input to the flow prediction neural network by forward warping the concatenation network output by. and; 【Number 18】 and; and; 【Number 19】 and; The method according to claim 6. **Claim 9** When the intersection region network is used to decode motion vectors, the decoding includes: receiving a residual latency and a motion vector latency from an encoder; combining the residual latency and the motion vector latency with a motion vector decoder network to generate a motion vector for use in motion compensation. and; and; generating the motion vector latency in the encoder includes using pixel values of the current picture and a previous reference picture as inputs to an optical flow network and a motion vector encoder network to generate the motion vector latency. and; The method according to claim 1. **Claim 10** When the intersection region network is used to decode a residual, the decoding includes: receiving a residual latency and a motion vector latency from an encoder; combining the residual latency and the motion vector latency with a residual decoder network to generate residual pixel values. and; generating the residual latency in the encoder includes: Accessing the residual pixel values of the current picture and the previous reference picture, accessing the motion vector latency, applying a motion vector decoder to the motion vector latency to generate a reconstructed motion vector, generating the residual latency based on the residual pixel values and the generated reconstructed motion vector comprising the method according to claim 1.

11. When entropy decoding uses spatio-temporal features, entropy decoding Estimate the spatial model parameters (φ t ) using the decoded neighboring latent features of the current picture to generate 【Number 20】 applying to a spatial context model, Hyper-prior parameter (Ψ t ) to estimate the hyper-prior decoding features 【Number 21】 applying to a hyper decoder, Temporal priority feature (γ t ) To estimate the latency characteristics of the previously decoded picture 【Number 22】 and applying the decoded motion from the current picture to a warping block, generating entropy model parameters for subsequent latency based on the spatial model parameters, the temporal prior features, and the hyper prior parameters having the method according to claim 1.

12. A method for improving the training of a neural network used in inter-frame coding, large motion training performed using random P-frame skips from 1 to n - 1 in the case of a training sequence with a total of n pictures, and a time-distance modulation loss that calculates a rate-distortion loss such as loss = w × lambda × MSE + Rate including one or more of Rate represents the achieved bitrate, MSE measures the distortion between the original picture and the corresponding reconstructed picture, and the weight parameter w 【Number 23】 is initialized based on the temporal frame distance as index i represents the number of iterations for N training iterations, method.

13. Calculating the modulation entropy loss Modulation entropy loss = (1 - p t ) b × (-log(p t )) having p t represents the estimated probability of the latency symbol t, b monotonically decreases from 5.0 to 0.0 over the N training iterations, the method according to claim 12.

14. A method for processing an uncompressed video frame with one or more neural networks, The motion vector information and spatial map information (α) of the non-compressed input video frame (x t ) are generated based on the continuity of the non-compressed input video frames including the non-compressed input video frame, and a motion-compensated frame [Number 24] generating based at least on a reference frame and a motion compensation network used to generate the motion vector information, applying the spatial map information to the motion-compensated frame to generate a weighted motion-compensated frame Generating a residual frame by subtracting the weighted motion-compensated frame from the uncompressed input video frame; Based on a residual encoder analysis network and a decoder synthesis network, a reconstructed residual frame 【Number 25】 is generated, and the residual encoder analysis network generates an encoded frame based on quantization of the residual frame; Generating a decoded approximation of the encoded frame by adding the weighted motion-compensated frame to the reconstructed residual frame A method having.

15. The spatial map information has weights within [0, 1], 0 indicates that only intra coding is prioritized, 1 indicates that only inter coding is prioritized, and weights between 0 and 1 represent intra-inter mixed coding. The method according to claim 14.

16. A non-transitory computer-readable storage medium storing computer-executable instructions for executing the method according to any one of claims 1 to 15 by one or more processors.

17. An apparatus having a processor and configured to execute the method according to any one of claims 1 to 15.