Tunable hybrid neural video representation

By introducing tunable hybrid implicit neural representation (T-NeRV) into video compression technology, combining specific frame embedding with GOP features, and using optical flow decoder, the problem of poor performance of the prior art in the fields of instability and high bit rate in dynamic sequences and high motion content processing is solved, achieving more efficient video compression effects.

CN120017841APending Publication Date: 2025-05-16DISNEY ENTERPRISES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411601081.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-18
Filing Date
2024-11-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing video compression techniques perform unstable when processing dynamic sequences and high motion content, and traditional methods perform poorly in the high bit rate field and are limited to specific aspect ratios and video sequences.

Method used

Tunable hybrid implicit neural representation (T-NeRV) solution is adopted, which enables video compression by combining specific frame embedding with features of specific groups of images (GOPs) and upsampling with optical flow-based decoders. This scheme captures embeddings through quantization perception and entropy constraint training, thereby jointly minimizing rate and distortion.

Benefits of technology

The T-NeRV solution outperforms the prior art in video representation and video compression tasks, and can effectively encode video frames under a variety of usage scenarios and parameters, improving performance for low motion sequences and dynamic sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017841A_ABST
    Figure CN120017841A_ABST
Patent Text Reader

Abstract

A system includes an adjustable neural network-based video encoder configured to receive a video sequence including a plurality of video frames, generate a particular frame embedding of a first video frame of the plurality of video frames, and identify one or more picture group (GOP) features of a subset of the plurality of video frames, the subset including the first video frame. The adjustable neural network-based video encoder is further configured to combine the particular frame embedding of the first video frame and the one or more GOP features of a first plurality of video frames of the plurality of video frames to provide potential features corresponding to a compressed version of the first video frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of and priority to pending U.S. provisional patent application serial number 63 / 599,972, filed on November 16, 2023, and entitled “TunableHybrid Neural Video Representations,” which is incorporated by reference into this application in its entirety. Background Art

[0003] Video compression is an old and difficult problem that has inspired a lot of research. The main goal of video compression is to represent digital video (usually a sequence of frames, each represented by a two-dimensional (2-D) pixel array, RGB or YUV colors) using the least amount of storage while minimizing quality loss. Although many advances have been made in traditional video codecs in recent decades, the emergence of deep learning has inspired many new methods that go beyond traditional video codecs.

[0004] For example, implicit neural representations (INRs) have attracted extensive research interest and have been applied to various fields including video compression. In addition to exhibiting desirable properties such as fast decoding and temporal interpolation capabilities, INR-based methods can match or exceed the compression performance of traditional standard video codecs such as Advanced Video Coding (AVC, also known as H.264) and High Efficiency Video Coding (HEVC). However, these existing methods that leverage INRs for video compression only perform well in limited and sometimes highly constrained settings, such as being restricted to specific model sizes, fixed aspect ratios, and relatively static video sequences.

[0005] As background, early video INRs were represented pixel-wise, mapping pixel indices to RGB colors, but with limited performance and poor decoding speed. However, a frame-by-frame representation approach, the Neural Representation of Video (NeRV), has been proposed. In NeRV, a multi-layer perceptron (MLP) generates temporal features from position-encoded frame indices that are reshaped and upsampled. NeRV achieves real-time decoding speed and significantly better reconstruction than previous video INRs. Inspired by the work on Generative Adversarial Networks (GANs), an accelerated neural network (E-NeRV) was developed, where the feature extractor used in traditional NeRV is decomposed into temporal and spatial contexts and fused via a Transformer Network, which further injects temporal information into the decoder block via an Adaptive Instance Normalization (AdaIN) layer.

[0006] Flow-guided frame-by-frame neural networks (FFNeRV) improve performance on dynamic sequences by enforcing temporal consistency using a decoder that predicts independent frames and a set of optical flow maps that are used to warp adjacent independent frames and combine independent frames to provide a final output frame. The hybrid representation HNeRV is similar to an autoencoder during training, using an encoder to extract content-specific embeddings. After training, the video is represented by the decoder's parameters and per-frame embeddings. While the existing variants of NeRV identified above, namely E-NeRV, FFNeRV, and HNeRV, represent improvements over the original NeRV, their usefulness is often limited to specific settings, such as performing well on static sequences but unstable on high-motion content, or vice versa. Traditional hybrid methods show potential in the low bitrate domain, but cannot scale to higher bitrates and are currently limited to videos with unusual 2:1 aspect ratios. Therefore, there is a need in the art for a neural network-based video compression solution that can efficiently encode video frames across a range of use cases and parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1A A diagram depicting an exemplary architecture of a video tunable hybrid implicit neural representation (T-NeRV) encoder for a video compression system according to one embodiment is shown;

[0008] Figure 1B A diagram depicting a device suitable for use according to one embodiment Figure 1A A diagram of an exemplary architecture of an optical flow based T-NeRV decoder of an exemplary T-NeRV encoder shown;

[0009] Figure 2 Schematically depicts a method for applying to a Figure 1B A series of T-NeRV blocks of a T-NeRV decoder as shown; and

[0010] Figure 3 is a flowchart outlining a method for performing tunable hybrid neural video representation according to one embodiment. DETAILED DESCRIPTION

[0011] The following description contains specific information related to the implementation in the present disclosure. Those skilled in the art will recognize that the present disclosure can be implemented in a manner different from that specifically discussed in this section. The drawings in this application and the detailed description attached thereto are only for exemplary embodiments. Unless otherwise specified, the same or corresponding elements in the drawings may be represented by the same or corresponding reference numerals. In addition, the drawings and illustrations in this application are usually not drawn to scale and are not intended to correspond to actual relative sizes.

[0012] The present application addresses and overcomes deficiencies in conventional techniques by using an adjustable hybrid implicit neural representation (INR) solution for video (hereinafter referred to as "T-NeRV"), which improves upon and develops the prior art based on the recognition that conventional codecs, such as Advanced Video Coding (AVC, also known as H.264) and High Efficiency Coding (HEVC), require both local and non-local information to efficiently encode frames. In addition, the adjustable hybrid neural video representation solution disclosed herein can be advantageously implemented as an automated or substantially automated system and method.

[0013] As used in this application, the terms "automation," "automated," and "automating" refer to systems and processes that do not require the involvement of a human user (e.g., a human system operator or administrator). Although in some embodiments, the performance of the systems and methods disclosed herein can be monitored by a human system operator, human involvement is optional. Therefore, the methods described in this application can be performed under the control of the hardware processing components of the disclosed systems.

[0014] The novel and inventive T-NeRV solution disclosed in this application implements a video compression system that includes a tunable neural network-based video encoder that combines specific frame embeddings with features of a specific group of images (specific GOPs) and upsamples them using an optical flow-based decoder. This novel and inventive combination of frame-specific and GOP-specific features provides leverage for content-specific fine-tuning. For example, larger frame embeddings, i.e., emphasizing frame-specific features, can extract more high-frequency information, thereby improving performance for low-motion sequences. Conversely, more significant GOP-specific features allow the T-NeRV model to extract more temporal context, which the decoder can use to reconstruct frames in dynamic sequences. For compression, the T-NeRV solution disclosed in this application captures embeddings through quantization-aware and entropy-constrained training to jointly minimize rate and distortion. By adopting a single entropy model for all embeddings, the present T-NeRV solution can exploit redundancy to a higher degree than previous methods. End-to-end training also forces the current T-NeRV network to fine-tune itself to the target video content by spending more bits on frame-specific embeddings or GOP-specific features, automatically adjusting the tuning lever.

[0015] The T-NeRV solution of the present invention contributes at least the following features to the state of the art: (i) an adjustable hybrid video INR that combines the embedding of a specific frame with the features of a specific GOP, thereby providing leverage for specific content fine-tuning; and (ii) an extension of the information-theoretic INR compression framework to include embeddings, thereby exploiting the significant redundancy therein. Note that during the training and optimization of the T-NeRV model, the adjustment of the model, as well as the balancing of the combination of the embedding of a specific frame with the features of a specific GOP, is performed automatically, although the system user can assert some control by limiting the maximum size of the feature maps used. It should also be noted that when evaluated on the UVG dataset, the T-NeRV solution disclosed in the present application outperforms all prior video INRs on both video representation and video compression tasks.

[0016] Figure 1A A diagram depicting an exemplary architecture of a T-NeRV encoder 110 of a video compression system according to one embodiment of the present concepts is shown, and Figure 1B A diagram depicting an exemplary architecture of an optical flow based T-NeRV decoder 140 suitable for use with an exemplary T-NeRV encoder 110 according to one embodiment is shown. The T-NeRV encoder 110 is configured to decode an optical flow-based T-NeRV decoder 140 according to a basic fact frame 102 (i.e., hereinafter “input video frame 102”) and normalized frame index 104(t) to generate latent features 118 (ie, ), the T-NeRV decoder 140 reconstructs the predicted frame 170 according to the latent features 118 (ie, ),like Figure 1B Note that the aspect ratio of the latent feature 118 matches the aspect ratio of the input video frame 102, i.e.

[0017] like Figure 1A As shown, the T-NeRV encoder 110 includes a large encoder 120, a position encoder (PE) 124, a multi-layer perceptron (MLP) 126, a multi-resolution feature grid 112, a fusion block 127, and in some embodiments, may also include an MLP 132. Figure 1A Also shown are learning weights 106, spatial frame embedding 122, one or more temporal features 128 (hereinafter referred to as "temporal features 128"), one or more specific GOP features 114 (hereinafter referred to as "GOP features 114"), a specific frame embedding 116 and an optional temporal feature output 130.

[0018] During training, the T-NeRV network, including the T-NeRV encoder 110 and the T-NeRV decoder 140, performs forward and backward passes, as known in the art, to optimize the network parameters. Once training is complete, only the GOP features 114, the specific frame embeddings 116 (i.e., the feature maps or tensor representations of the video frames), the latent features 118, and the optional temporal feature outputs 130 are retained as part of the video state, while Figure 1A The remaining features shown in are discarded. Therefore, decoding the input video frame 102 is equivalent to obtaining GOP features 114 from the multi-resolution feature grid 112 and fusing them with the specific frame embedding 116 using any suitable combination technique (such as concatenation or fusion) to obtain the latent features 118 (z t ), the latent feature is passed through the T-NeRV decoder 140.

[0019] Specific frame embedding 116 can be one of the two components of the latent features 118. The specific frame embedding 116 can be generated in two steps: first, the large encoder 120 can extract the spatial frame embedding 122 from the input video frame 102 It can then be augmented with additional temporal features 128. This augmentation is a mathematical operation that combines the spatial frame embedding 122 and the temporal features 128 by concatenating them, or alternatively, combining them using any other mathematical technique, such as a neural network or a transformer. Because the specific frame embeddings themselves are transmitted to the decoder, all network blocks involved in generating them can be discarded after encoding. As a result, the large encoder 120 can be approximately 100 times larger than a conventional encoder used in a HNeRV network, i.e., include 100 times more weights, because the encoder weights themselves are not transmitted to the decoder.

[0020] The spatial frame embedding 122 may be further augmented with temporal features 128. For example, a position encoder (PE) 124 followed by a multi-layer perceptron (MLP) 126 may be used to provide temporal information that may then be reshaped to match the spatial frame embedding 122. The time characteristics of the dimension 128 As described above, the temporal features 128 can then be embedded with the spatial frame 122 Combined to obtain a specific frame embedding 116, which can be expressed as:

[0021]

[0022] where α t represents the learned weights 106 specific to each frame, which modulate the temporal features 128 (s t) is included in the particular frame embedding 116. This allows the T-NeRV network to decide on a per-frame basis whether and to what extent to incorporate the temporal features 128 into the particular frame embedding 116. Note that the mechanism used to combine the spatial frame embedding 122 and the temporal features 128 to provide the particular frame embedding 116 can be considered a per-frame masking operation on the particular frame embedding 116, so that similar frames obtain different embeddings. The T-NeRV decoder 140 can use this information to decide which parts of the frame to reconstruct from independent frames of the frame itself, and which parts rely on inter-frame information.

[0023] The T-NeRV encoder 110 uses a multi-resolution feature grid 112 to obtain GOP features 114 During the forward pass, the frame index 104(t) is used to index the grid, and two features are selected at each level of the multi-resolution feature grid 112, which are in the form of learned numbers or tensors, and t is located within these features. Bilinear interpolation can be performed between these two features, and the results from each level can be concatenated, fused, or otherwise combined to obtain the GOP features 114. Note that in contrast to FFNeRV, T-NeRV allows features at different levels of the multi-resolution feature grid 112 to be compared with respect to channel c. i The number of is varied in size while maintaining their spatial dimensions (h, w). This modification allows T-NeRV to advantageously adjust the amount of information captured for different GOP sizes. Note that this adjustment of T-NeRV occurs automatically during the optimization performed during the learning process.

[0024] The latent features 118 obtained from the fusion of GOP features 114 and specific frame embedding 116 can be combined with the temporal feature output 130 generated by the MLP 132. are passed to the T-NeRV decoder 140. Figure 1B In the exemplary embodiment shown, the T-NeRV decoder 140 may include three parts: a pre-block convolution layer 142, a series of T-NeRV blocks 150 (e.g., five T-NeRV blocks), and a flow-based warp block 144. Note that the pre-block convolution layer 142 includes a convolution layer with a kernel that performs channel expansion or channel reduction, i.e., a tensor reshaping operation known in the art, depending on the size of the latent features 118 and the size expected by the first T-NeRV block 150. Figure 1B Also shown are the head layers, namely output layers 164a and 164b, optical flow map 143, independent frame buffer 146, aggregated frame 148, independent frame 166, aggregated block 168, and corresponding Figure 1A The predicted frame 170 of the input video frame 102 is shown in FIG.

[0025] The T-NeRV decoder 140 may utilize a series of T-NeRV blocks 150 to predict independent frames 166 and optical flow maps 143. Figure 2 , according to one embodiment, Figure 2 Schematically depicts the Figure 1B A series of T-NeRV blocks 250 of the T-NeRV decoder 140 in FIG. Figure 2 As shown, the linear layer 252 extracts per-channel statistics such as mean and variance from the temporal feature output 230 from the T-NeRV encoder 110, which are used to shift the distribution of the input tensor through the adaptive instance normalization (AdaIN) layer 254, thereby injecting temporal information into the series of T-NeRV blocks. Note that the temporal feature output 230 and the series of T-NeRV blocks 250 generally correspond to the temporal feature output 130 and the series of T-NeRV blocks 150, respectively. Figure 1A and 1B Thus, the temporal feature output 130 and the series T-NeRV block 150 may share any features attributed by the present disclosure to the corresponding temporal feature output 230 and the series T-NeRV block 250, and vice versa.

[0026] like Figure 2 As further shown in , in addition to the linear layer 252 and the AdaIN layer 254, the first two T-NeRV blocks of the series of T-NeRV blocks 150 / 250 include a convolution layer 256 with channel expansion (e.g., a 5×5 convolution layer) and a subsequent pixel shuffle (PixelShuffle) layer 258, which can perform spatial upsampling and can be followed by an activation function 260a in the form of a Gaussian Error Linear Unit (GELU) activation function. In some embodiments, the T-NeRV decoder 140 can utilize a low channel reduction factor, such as r=1.2, and increase the kernel size. This approach allows the T-NeRV decoder 140 to allocate more parameters to later stages, facilitating the reconstruction of high-frequency details. Specifically, this strategy makes the third T-NeRV block of the series of T-NeRV blocks 150 / 250 the largest, which is advantageous given its dual purpose in upsampling and optical flow prediction.

[0027] Note that although HNeRV teaches that kernel size k i =min{2i-1,5}, but it is found that the 5×5 convolution is computationally expensive and parameter-inefficient. In further contrast to HNeRV, the third to fifth T-NeRV blocks of the series of T-NeRV blocks 150 / 250 use two consecutive convolutional layers 262 (e.g., 3×3 convolutional layers) with the same total number of parameters and additional activations 260b between the convolutional layers 262, resulting in Figure 2The higher parameter efficiency and the additional nonlinearity further help the T-NeRV decoder 140 to reconstruct the high frequency information.

[0028] exist Figure 1B In the example, the flow-based warping block 144 acts as a flow-guided aggregation module for reusing information from neighboring frames via optical flow. To this end, the head layer 164a, i.e., one or more output layers after the fifth T-NeRV block of the series of T-NeRV blocks 150 / 250, outputs an independent frame 166 and two aggregation weights w I Copy independent frame 166 (I t ) is separated from the computation graph and stored in a separate frame buffer 146. The aggregation window can then be defined For example to utilize information from the previous two frames and the next two frames. Figure 1A In the forward pass of frame index 104, t, after the third T-NeRV block of the series of T-NeRV blocks 150 / 250, the header layer 164b is for each aggregation window j predicted optical flow map 143 and weight graph The optical flow map and weight map can then be bilinearly upsampled and normalized by a softmax function. Then, the flow-based warping block 144 can use the optical flow map M(t, t+j) to warp the corresponding adjacent independent frames I t+j and the results may be aggregated in a weighted manner to obtain an aggregate frame 148A t ,as follows:

[0029]

[0030] Finally, in another weighted aggregation at aggregation block 168, the aggregate frames 148 (A t ) and independent frame 166 (I t ) to obtain the predicted frame 170, that is, the predicted frame 170

[0031] Regarding the training of the T-NeRV network including the T-NeRV encoder 110 and the T-NeRV decoder 140, it is noted that INR compression can be modeled as a rate-distortion problem L=D+λR, where D represents some distortion loss, R represents the entropy of the INR parameter θ, and λ establishes a balance between the two. The training INR loss L has the desired result of minimizing both rate and distortion during training, thereby allowing them to achieve compression by entropy coding the parameter θ after training. Therefore, the T-NeRV network including the T-NeRV encoder 110 and the T-NeRV decoder 140 can be trained using an optimization loss function including a distortion loss and an entropy loss, where the distortion loss and the entropy loss can be jointly optimized.

[0032] In addition, because the above training process requires a discrete set of symbols, such as the set of integers So we perform quantization-aware training using a straight-through estimator (STE) known in the art to ensure differentiability. We quantize each layer independently using scale reparameterization with two trainable parameters. Then, we can estimate the entropy of each layer independently by fitting a small neural network to the weight distribution.

[0033] The same framework is then extended to handle parameter encoding, such as the multi-resolution feature grid 112 of the T-NeRV encoder 110, and embedding. Each of the multiple feature grids included in the multi-resolution feature grid 112, such as three feature grids, is treated as its own layer with its own quantization parameters. During the forward pass, each feature grid is quantized and dequantized using the same scale reparameterization scheme before bilinear interpolation of the two features. Note that employing an entropy model for each feature grid allows the T-NeRV network to determine the degree to which the features of each temporal frequency should be compressed.

[0034] In contrast to the approach described above with reference to the multi-resolution feature grid 112, a single entropy model is employed for all frame-specific embeddings 116, thereby encouraging the T-NeRV network to exploit redundancy between frames. This balances the injection of temporal features 128 into the frame-specific embeddings 116 described above, constraining the T-NeRV network to use such temporal information only when the gain in video quality outweighs the increase in entropy. Like the distortion loss, the entropy loss is back-propagated through the T-NeRV network, which encourages the T-NeRV network to learn features in its feature grid that exhibit low entropy.

[0035] Will refer to Figure 3 Further description Figure 1A The functionality of the T-NeRV encoder 110 is shown. Figure 3 A flowchart 380 is shown outlining an exemplary method for performing a tunable hybrid neural video representation according to one embodiment. Figure 3380, noting that certain details and features have been omitted from flowchart 380 so as not to obscure the discussion of the inventive features of the present application.

[0036] Combination Figure 1A refer to Figure 3 , the method outlined in flowchart 380 includes receiving a video sequence including a plurality of video frames (eg, input video frame 102) by T-NeRV encoder 110 (action 381). Figure 1A As shown, the T-NeRV encoder 110 may use the large encoder 120 of the T-NeRV encoder 110 to receive the input video frame 102 and other video frames included in the video sequence received in action 381. As described above, the large encoder 120 may be approximately one hundred times larger, i.e., include one hundred times more weights, than a traditional encoder used in a HNeRV network.

[0037] Continue to combine references Figure 3 and 1A , the method outlined by flowchart 380 also includes generating a specific frame embedding 116 of a first video frame (hereinafter referred to as "input video frame 102") of the video frames included in the video sequence received in act 381 (act 382). As described above, the specific frame embedding 116 can be generated in one or two steps: first, the large encoder 120 extracts the spatial frame embedding 122 from the input video frame 102, which can then be used as the specific frame embedding 116 in some use cases, or can be augmented with additional temporal features 128 in other use cases.

[0038] The spatial frame output by the large encoder 120 is embedded in 122 In embodiments further augmented with temporal features 128 to provide a specific frame embedding 116, PE 124 and subsequent MLP 126 may be utilized to extract temporal information from frame index 104. The temporal information may be extracted from a subset of video frames included in a video sequence received in act 381, wherein the subset of video frames of the video sequence includes input video frame 102. The temporal information output by MLP 126 may then be reshaped into temporal features 128. The dimensions of the spatial frame embedding 122 are matched. The temporal features 128 may then be fused into the spatial frame embedding 122 to obtain the specific frame embedding 116. Thus, in action 382, ​​the T-NeRV encoder 110 may use the large encoder 120 to provide the spatial frame embedding 122, and in some use cases, combine the spatial frame embedding 122 with the temporal features 128 to generate the specific frame embedding 116.

[0039] Continue to combine references Figure 3 and 1A, the method outlined in flowchart 380 also includes identifying one or more GOP features of a subset of video frames included in the video sequence received in action 381, wherein the subset of video frames includes the input video frame 102 (action 383). Such GOP features may include, for example, I frames (intra-coded pictures), P frames (predicted pictures), and B frames (bidirectional pictures), as well as integers M and N, wherein M represents the number of frames between two anchor frames, i.e., the distance between an I frame and a P frame or the distance between two P frames, and N represents the number of frames between two I frames. Note that the subset of video frames from which the one or more GOP features are identified in action 383 may be the same subset of video frames from which the temporal features 128 are optionally extracted as part of action 382, ​​or another subset of video frames of the video sequence received in action 381 that also includes the input video frame 102.

[0040] Action 383 may be performed by the T-NeRV encoder 110 using the multi-resolution feature grid 112 to obtain one or more GOP features. As described above, during the forward pass, the frame index 104(t) is used to index into the grid, and at each level of the multi-resolution feature grid 112 two features are selected where the frame index 104 is located. Bilinear interpolation may be performed between the two features, and the results from each level may be concatenated, fused, or otherwise combined to obtain the GOP feature 114. As described above, in contrast to FFNeRV, T-NeRV allows features at different levels of the multi-resolution feature grid 112 to be compared with respect to the channel c. i The number of φ varies in size while maintaining their spatial dimensions (h, w). This modification allows T-NeRV to advantageously adjust the amount of information captured for different GOP sizes.

[0041] Combination Figure 3 and Figure 1A, the method outlined in flowchart 380 further includes combining the specific frame embedding 116 of the input video frame 102 with the GOP features 114 identified in action 383, either by concatenating them, or combining them using any other mathematical method, neural network or transformer, to provide a latent feature 118 corresponding to a compressed version of the input video frame 102 (action 384). As described above, the T-NeRV encoder 110 is a tunable neural network based video encoder configured to combine the specific frame embedding 116 with the group of GOP features using a fusion block 127 to concatenate them, or alternatively, combine them using, for example, any other mathematical technique, neural network or transformer. The combination of the specific frame embedding 116 with the GOP features 114 provides leverage for specific content fine-tuning, because emphasizing the specific frame embedding 116 extracts more high-frequency information from the video frame, thereby improving performance for low motion sequences. Conversely, giving more prominence to the GOP features 114 allows the T-NeRV model to extract more temporal context, which the decoder 140 can use to better reconstruct frames in dynamic sequences. As described above, this leveraging, or balancing of the combination of specific frame embeddings with features for a specific GOP, is performed automatically during training and optimization of the T-NeRV model, although the system user can assert some control by limiting the maximum size of the feature maps used.

[0042] Thus, in action 384, the T-NeRV encoder 110 may provide the latent feature 118 as a weighted or unweighted combination of the specific frame embedding 116 and the GOP feature 114. Furthermore, in some embodiments where the latent feature 118 is provided as a weighted combination of the specific frame embedding 116 and the GOP feature 114, one or both of the first weight applied to the specific frame embedding 116 or the second weight applied to the GOP feature in the weighted combination may be determined by the T-NeRV encoder 110 based on, for example, whether the content depicted in the video sequence received in action 381 is primarily static or primarily dynamic.

[0043] Combined with reference Figure 1A , 1B and Figure 3 In some embodiments, the method outlined in flowchart 380 may further include outputting the latent features 118 of the input video frame 102 to a decoder, such as the T-NeRV decoder 140, or in other embodiments, even another type of conventional NeRV-based decoder (action 385). Note that action 385 is optional, and in some embodiments, the method outlined in flowchart 380 may end with action 384 described above. Optional action 385, when executed, is performed by the T-NeRV encoder 110, and, as described above, may also include providing the temporal feature output 130 to the decoder.

[0044] Note that although actions 382, ​​383, and 384 or actions 382, ​​383, 384, and 385 are described above as being performed on a single video frame (e.g., input video frame 102) at a time, in some embodiments, these actions may be performed on segments of a video sequence (hereinafter referred to as a "subsequence"). That is, in some use cases, input video frame 102 may correspond to a subsequence of a subset of video frames identified in action 383, specific frame embedding 116 may include multiple specific frame embeddings corresponding to each video frame of the subsequence of the identified subset of video frames, and latent feature 118 may be multiple latent features, each of which corresponds to a compressed version of a video frame of the subsequence of the identified subset of video frames. In embodiments where latent feature 118 corresponds to multiple latent features, in action 385, these multiple latent features may be output to a decoder, or may be stored for later decoding.

[0045] In some embodiments, the method outlined in flowchart 380 may end with action 385 as described above. However, in embodiments where the video compression system including the T-NeRV encoder 110 also includes a T-NeRV decoder 140, the method outlined in flowchart 380 may also include actions of receiving, by the T-NeRV decoder 140, the latent features 118 from the T-NeRV encoder 110, and decoding, by the T-NeRV decoder 140, the latent features 118 to provide a predicted frame 170 as an uncompressed video frame corresponding to the input video frame 102. Note that the method may be performed by the T-NeRV decoder as described above with reference to Figure 1B and Figure 2 The decoding of the latent features 118 into the provided prediction frame 170 is performed in a manner described in detail.

[0046] With respect to the method outlined by flowchart 380 and described above, it is noted that actions 381 , 382 , 383 , and 384 , or actions 381 , 382 , 383 , 384 , and 385 , may be performed in an automated process, wherein human involvement may be omitted.

[0047] Thus, the present application discloses systems and methods for performing adjustable hybrid neural video representations, which advance the prior art by at least the following contributions: (i) developing a T-NeRV encoder that combines specific frame embeddings with features of specific GOPs to provide leverage for specific content fine-tuning; and (ii) extending the information-theoretic INR compression framework to include embeddings, thereby advantageously exploiting the significant redundancy therein. As described above, when evaluated on the UVG dataset, the T-NeRV solution disclosed in the present application outperforms all prior video INRs on both video representation and video compression tasks.

[0048] From the above description, it is apparent that various techniques may be used to implement these concepts without departing from the scope of the concepts described in this application. In addition, although these concepts have been described with specific reference to certain embodiments, it will be appreciated by those of ordinary skill in the art that changes may be made in form and detail without departing from the scope of these concepts. Therefore, the described embodiments are considered to be illustrative and non-restrictive in all respects. It should also be understood that the application is not limited to the specific embodiments described herein, but that many rearrangements, modifications, and substitutions are possible without departing from the scope of this disclosure.

Claims

1. A system comprising: The video encoder based on a tunable neural network is configured as follows: receiving a video sequence comprising a plurality of video frames; generating a specific frame embedding of a first video frame of a plurality of video frames; identifying one or more group of pictures (GOP) characteristics of a first plurality of video frames of a plurality of video frames, the first plurality of the plurality of video frames including the first video frame; and The particular frame embedding of the first video frame and one or more GOP features of a first plurality of video frames of the plurality of video frames are combined to provide a latent feature corresponding to a compressed version of the first video frame.

2. The system of claim 1, wherein one or more GOP features of a first plurality of video frames of the plurality of video frames are identified using a multi-resolution feature grid of a tunable neural network based video encoder.

3. The system of claim 1, wherein the tunable neural network based video encoder is configured as: extracting one or more temporal features of a first plurality of video frames of the plurality of video frames from a first plurality of video frames of the plurality of video frames; Wherein a specific frame embedding is generated based on a combination with one or more temporal features.

4. The system of claim 1, wherein the latent features include a weighted combination of a specific frame embedding and one or more GOP features.

5. The system of claim 4, wherein at least one of the first weight applied to the embedding of a particular frame or the second weight applied to one or more GOP features is determined by a tunable neural network based video encoder.

6. The system according to claim 1, wherein: The tunable neural network based video encoder is trained using an optimization loss function including a distortion loss and an entropy loss. The system of claim 6 , wherein the distortion loss and the entropy loss are jointly optimized.

8. The system of claim 1 , wherein the first video frame comprises a subsequence of a first plurality of video frames of a plurality of video frames, and wherein the specific frame embedding comprises a plurality of specific frame embeddings corresponding respectively to each subsequence of the first plurality of video frames of the plurality of video frames, and wherein the latent feature is one of a plurality of latent features, each latent feature corresponding respectively to a compressed version of each subsequence of the first plurality of video frames of the plurality of video frames.

9. A system comprising: The neural network based video decoder is configured as follows: receiving a latent feature of a compressed version of a first video frame of a first plurality of video frames corresponding to a plurality of video frames included in a video sequence, the latent feature being a combination of a particular frame embedding of the first video frame and one or more group of pictures (GOP) features of the first plurality of video frames of the plurality of video frames; and The latent features are decoded to provide an uncompressed video frame corresponding to the first video frame.

10. The system according to claim 9, wherein: The specific frame embedding is generated based on a combination with one or more temporal features extracted from a first plurality of video frames of the plurality of video frames.

11. The system of claim 9, wherein the latent features include a weighted combination of a specific frame embedding and one or more GOP features.

12. The system of claim 9, wherein the first video frame comprises a subsequence of a first plurality of video frames of the plurality of video frames, and wherein the specific frame embedding comprises a plurality of specific frame embeddings corresponding respectively to each subsequence of the first plurality of video frames of the plurality of video frames, and wherein the latent feature is one of a plurality of latent features, each latent feature corresponding respectively to a compressed version of each subsequence of the first plurality of video frames of the plurality of video frames.

13. A method performed by a video encoder based on a tunable neural network, the method comprising: receiving a video sequence comprising a plurality of video frames; generating a specific frame embedding of a first video frame of a plurality of video frames; identifying one or more group of pictures (GOP) characteristics of a first plurality of video frames of the plurality of video frames, the first plurality of video frames of the plurality of video frames including the first video frame; and The particular frame embedding of the first video frame and one or more GOP features of a first plurality of video frames of the plurality of video frames are combined to provide a latent feature corresponding to a compressed version of the first video frame.

14. The method of claim 13, wherein one or more GOP features of a first plurality of video frames of the plurality of video frames are identified using a multi-resolution feature grid of the tunable neural network based video encoder.

15. The method according to claim 13, further comprising: extracting one or more temporal features of a first plurality of video frames of the plurality of video frames from a first plurality of video frames of the plurality of video frames; Wherein a specific frame embedding is generated based on a combination with one or more temporal features.

16. The method of claim 13, wherein the latent features include a weighted combination of a specific frame embedding and one or more GOP features.

17. The method according to claim 16, further comprising: At least one of a first weight applied to a particular frame embedding or a second weight applied to one or more GOP features is determined.

18. The method according to claim 13, wherein: The tunable neural network based video encoder is trained using an optimization loss function including a distortion loss and an entropy loss. The method of claim 18 , wherein the distortion loss and the entropy loss are jointly optimized.

20. The method of claim 13, wherein the first video frame comprises a subsequence of a first plurality of video frames of a plurality of video frames, and wherein the specific frame embedding comprises a plurality of specific frame embeddings corresponding respectively to each subsequence of the first plurality of video frames of the plurality of video frames, and wherein the latent feature is one of a plurality of latent features, each latent feature corresponding respectively to a compressed version of each subsequence of the first plurality of video frames of the plurality of video frames.