Video neural representation and generalization end-to-end coding and decoding combined video coding and decoding method and device

By combining end-to-end and video neural representation methods with low-rank adaptation, the approach addresses the challenge of adapting generic knowledge to specific video content, enhancing compression performance and decoding efficiency.

CN120321396APending Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410056316.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing video encoding and decoding methods based on deep learning, especially the generalized end-to-end method, have the problem that prior knowledge does not exactly match the actual compressed video, resulting in poor restoration of video frames. The existing online learning technology only optimizes the encoding end model but fails to fully improve the performance of the decoding end model.

Method used

Combining the generalized end-to-end encoding and decoding method and video neural characterization, by introducing an efficient fine-tuning model to fine-tune the model at the decoding end, and full parameter optimization is carried out on the encoding end, and the parameter transmission cost is reduced by using technologies such as LoRA to achieve enhanced adaptability to the decoding end model.

Benefits of technology

It improves the compression performance of video encoding and decoding methods, enhances the adaptability to specific content, reduces the cost of video storage and transmission, and can further improve performance in combination with existing online learning technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321396A_ABST
    Figure CN120321396A_ABST
Patent Text Reader

Abstract

The invention discloses a video nerve representation and generalization end-to-end coding and decoding combined video coding method, decoding method and device, and the method comprises the steps: 1) at a coding end, adding and creating an efficient fine tuning model in a decoding end model according to a network structure of the generalization end-to-end video coding and decoding method; taking a video to be coded as the input of an optimization process, optimizing the efficient fine tuning model and generalizing a coding end model of the end-to-end video coding and decoding method; the coding end obtains potential features through the optimized model, quantifies the potential features and performs entropy coding to form a code stream; the coding end quantizes the parameters of the efficient fine tuning model and performs entropy coding to form a code stream; 2) at a decoding end, receiving the potential feature code stream and the code stream of the efficient fine tuning model parameters, and restoring the code stream into the efficient fine tuning model and the potential features; and decoding the potential features into a video by using a decoding end model of a generalization end-to-end coding and decoding method combined with an efficient fine tuning model. The method has the advantages that the adaptability of the coding end and the decoding end of the generalization end-to-end method to specific contents can be enhanced at the same time, and the method can be used for enhancing the performance of most generalization end-to-end coding and decoding methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video encoding and decoding, and particularly to end-to-end intelligent video encoding and decoding technology. Background Art

[0002] Video is a very important information medium. With the continuous development of video services, the amount of video data has been growing rapidly. Users' increasing pursuit of video frame rate and clarity poses more challenges to existing video encoding and decoding methods. How to further improve video encoding and decoding performance is an important issue.

[0003] In the past decade, deep learning-based methods have made great progress in the field of video encoding and decoding. Different from traditional video encoding and decoding methods, these methods use deep neural network models to replace each module in traditional video encoding and decoding methods. The function of each neural network model is jointly determined by the model structure and internal parameters. The model structure is defined by engineers in advance, and the internal parameters are continuously optimized through methods such as loss calculation and gradient descent (i.e., the training method) until the model can encode and decode video well. When the neural network model is trained, these models can be used to implement video encoding and decoding. Deep learning-based methods generally have two major steps. The first step is to train the neural network model, and the second step is to use these neural network models for video encoding and decoding.

[0004] There are mainly two existing video encoding methods completely based on deep learning. One is the generalized end-to-end method, and the other is the overfitting method based on video neural representation. The following will be combined with Figure 1 introduced separately. Figure 1 In (A) is the generalized end-to-end video encoding and decoding method Figure 1 (B) is the video encoding and decoding method based on video neural representation.

[0005] Among the two methods, the dominant one is the generalizable method, that is Figure 1 Method A in, for example: DVC (DVC: An End-to-end Deep Video Compression Framework, end-to-end deep video compression framework), DCVC (Deep Contextual Video Compression, deep contextual video compression). Generalizable means that after training all neural network models on a large-scale dataset (this process is generally called the pre-training process), the existing models can be used to encode and decode any video without repeated training. End-to-end means that the input video is completely processed by the neural network model, and the decoded video is also completely output by the neural network model, and all models can be trained together.

[0006] Models of the generalized end-to-end method generally include: an encoder model, a decoder model, an entropy coding model, a context mining model, etc. They can be divided into an encoding-end model and a decoding-end model according to whether they are used at the encoding end or the decoding end. These models all need to be pre-trained. During pre-training, the video encoding and decoding process is continuously simulated, so that the model can learn knowledge related to video encoding and decoding through the feedback of the objective function defined by the engineer in each simulation. The embodiment of this knowledge is the parameters inside each model. After pre-training, actual video encoding and decoding can be performed. At this time, the videos for encoding and decoding are generally videos not seen during pre-training, and the model can only use the knowledge learned during pre-training for encoding and decoding. The process of video encoding and decoding using the pre-trained neural network model is as follows:

[0007] 1) The encoding end (i.e., the owner of the uncompressed video) uses the pre-trained encoder to convert each frame image of the original video into latent features. The latent features are more compact than the original video frames and are more conducive to compression.

[0008] 2) The encoding end uses the pre-trained entropy coding model to analyze the latent features and converts them into a binary bitstream through quantization and entropy coding for transmission to the decoding end.

[0009] 3) After receiving the bitstream of the latent features, the decoding end (i.e., the receiver of the compressed video) uses the pre-trained entropy coding model to restore it to the latent features, and then uses the pre-trained decoder model to convert the latent features into reconstructed video frames.

[0010] The generalized method hopes to obtain certain prior knowledge from a large amount of data and then use this prior knowledge to compress any video. However, the prior knowledge obtained from the large-scale dataset may not be fully applicable to the actual compressed videos. For these videos, the model may not be able to restore the video frames well, and there may also be a large deviation in the probability estimation of the latent features, which means that the pre-trained model still has room for further improvement.

[0011] The second method, the video encoding and decoding method based on video neural representation, that is Figure 1Method B in [reference], this method does not pre-train on large-scale datasets, but directly trains for a specific video to obtain the feature set and decoder model of that video or a separate decoder model. If there is only a separate decoder model, the input of the decoder model is generally a value agreed upon between the encoder and the decoder. For example, NeRV (NeRV: Neural Representations for Videos) designed a convolutional network that predicts the image corresponding to a time index based on the time index, which is to train a separate decoder model for a specific video. Just send the parameters of this convolutional network to the decoder, and the user can restore the decoder according to the agreed decoder structure. At the same time, input the desired video frame index, and the image of that frame can be obtained. Another example is FFNeRV (FFNeRV: Flow-Guided Frame-Wise Neural Representations for Videos), which designed a feature set and a convolutional network that predicts an image based on latent features, which is to train the feature set and decoder model of a specific video for that video. The feature set also needs to be transmitted to the decoder to restore the video.

[0012] Compared with the generalization method, these methods not only convert video information into latent features, but also convert a part of the information into the parameters of the neural network model (that is, the parameters of the decoder model). The models of these methods are simple, but they are still in the early stage of research. The current compression performance is not yet sufficient to be compared with the state-of-the-art generalization methods. The main reason may be that there is no knowledge sharing between the networks representing different videos, resulting in overall parameter redundancy, and there is no mature scheme for jointly optimizing the entropy coding and decoder model. However, this method provides an idea of converting video information into neural network parameters.

[0013] The above two methods each have their own advantages and problems. However, due to performance differences and the issue of development sequence, the current mainstream deep learning-based video coding and decoding method is still the generalization end-to-end video coding and decoding method. Aiming at the problem that the prior knowledge of the generalization method does not fully match the actual compressed video, online learning technology proposes to perform online optimization on the model of the generalization coding and decoding method, such as a work published in ECCV (European Conference on Computer Vision) in 2020: Content Adaptive and Error Propagation Aware Deep Video Compression. Such as Figure 2As shown, the online learning method does not add or delete the model of the generalization encoding and decoding method. Instead, after the generalization method is pre-trained, the encoding end model is fine-tuned according to the specific content of the encoded video, or the latent features of the video are directly fine-tuned. This method enables the model to make appropriate adjustments according to the content, and can enhance the performance of the generalization video encoding and decoding method.

[0014] However, this technology only optimizes the model at the encoding end, while the model at the decoding end remains unchanged. Compared with the fully optimizable scheme, the performance improvement space of this technology is relatively limited. Summary of the Invention

[0015] To overcome the above defects of the prior art, inspired by video neural representations, the present invention proposes a video encoding and decoding method that combines a generalization end-to-end encoding and decoding method with video neural representations to solve the problem that the prior knowledge of the generalization method does not fully match the specific video in a new way. The specific concept of the present invention is as follows:

[0016] In order to better compress the specific video content by the generalization end-to-end video encoding and decoding method, the present invention needs to fine-tune the general encoding and decoding model. Existing methods are limited by the idea of the generalization end-to-end encoding and decoding method, only transmitting the latent features of the video, and not considering sending a new decoding model to the decoding end. Therefore, the decoding end model cannot be optimized. The present invention draws on the idea of video neural representations, which can transmit neural network parameters as video information and update the decoding end model by transmitting the neural network model parameters. However, if all models are directly fine-tuned with full parameters, the parameters of the decoding end model will change greatly, and all parameters of the decoding end model need to be transmitted, and the cost of transmitting parameters will be very high.

[0017] Therefore, the present invention adopts an efficient fine-tuning method in the field of large language models, such as LoRA (LoRA: Low-Rank Adaptation of Language Models), etc., to introduce an additional highly efficient fine-tuning model with very low parameter quantity to fine-tune the decoding end model without modifying the parameter of the decoding end model of the generalization end-to-end video encoding and decoding method, saving the cost of parameter transmission. In addition, since the encoding end model is not used at the decoding end and does not need to be transmitted, the present invention uses the full parameter fine-tuning method to optimize the encoding end model, and the two are combined to achieve the adaptation of the entire encoding and decoding model to the compressed content.

[0018] The first object of the present invention is to propose a video encoding method that combines video neural representations with generalization end-to-end encoding and decoding, and the encoding method includes the following steps:

[0019] 1) According to the network structure of the generalization end-to-end video encoding and decoding method, add and create an efficient fine-tuning model in its decoding end model;

[0020] 2) Use the video to be encoded as the input of the optimization process, and optimize the encoding-end model of the efficient fine-tuning model and the end-to-end video encoding method for generalization;

[0021] 3) The encoding end obtains latent features through the optimized model, and quantizes and entropy-codes the latent features into a bitstream;

[0022] 4) The encoding end quantizes and entropy-codes the parameters of the efficient fine-tuning model into a bitstream.

[0023] The second object of the present invention is to propose a video decoding method combining video neural representation and end-to-end coding and decoding for generalization, which includes the following steps:

[0024] 1) The decoding end receives the bitstream of latent features and the bitstream of the parameters of the efficient fine-tuning model, and restores the bitstreams to the efficient fine-tuning model and latent features;

[0025] 2) The decoding end uses the decoding-end model that combines the end-to-end method for generalization with the efficient fine-tuning model to decode the latent features into a video.

[0026] The third object of the present invention is to provide a video encoding device combining video neural representation and end-to-end coding and decoding for generalization, which is characterized by including:

[0027] A processor;

[0028] A memory for storing the video to be encoded and the parameters of the neural network model;

[0029] A signal transmitter;

[0030] And several programs running on the processor and the signal transmitter:

[0031] Model optimization program: The processor uses this program to read the video to be encoded and optimize the encoding-end model of the end-to-end method for generalization and the efficient fine-tuning model with this video;

[0032] Encoding program: The processor uses this program to quantize and entropy-code the latent features and the parameters of the efficient fine-tuning model into a bitstream and send it to the signal sending device;

[0033] Signal sending program: The signal transmitter uses this program to send the bitstream to the receiving device.

[0034] The fourth object of the present invention is to provide a video encoding device combining video neural representation and end-to-end coding and decoding for generalization, which is characterized by including:

[0035] A processor;

[0036] A memory for storing the decoded video and the parameters of the neural network model;

[0037] Signal receiver;

[0038] And several programs running on the processor and signal receiver:

[0039] Signal receiving program: The receiver receives signals and transmits the code stream to the processor;

[0040] Decoding program: The processor uses this program to restore the code stream to an efficient fine-tuning model and latent features. The processor uses the decoding-end model and the efficient fine-tuning model to decode the latent features into video.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] (1) The present invention is a video coding and decoding method that combines the generalized end-to-end coding and decoding method and video neural representation. Based on the generalized end-to-end coding and decoding model, a video neural representation for specific content is added, which can enhance the adaptability of the codec to this content and can enhance the compression performance of most generalized end-to-end video coding and decoding methods.

[0043] (2) The present invention can be combined with existing online learning technologies to enhance the performance of existing generalized end-to-end video coding and decoding methods. Because the video latent features used for transmitting information in the present invention can be optimized frame by frame using online learning methods.

[0044] (3) The video neural representation obtained by the present invention can be regarded as a highly generalized form of video information, which helps to reduce the cost of video storage and transmission. Description of the Drawings

[0045] Figure 1 Comparison diagram of the prior art generalized end-to-end video coding and decoding method and the video coding and decoding method based on video neural representation;

[0046] Figure 2 Schematic diagram of the prior art online learning enhanced generalized end-to-end video coding and decoding method;

[0047] Figure 3 Schematic diagram of the principle of the coding and decoding method of the embodiment of the present invention;

[0048] Figure 4 Schematic diagram of the flow of the coding method of the embodiment of the present invention;

[0049] Figure 5 Flowchart of adding and creating an efficient fine-tuning module in the decoding-end model according to the network structure of the generalized end-to-end video coding and decoding method of the embodiment of the present invention;

[0050] Figure 6 Schematic diagram of the Conv-Adapter fine-tuning convolutional network method of the embodiment of the present invention;

[0051] Figure 7 Schematic diagram of the fine-tuning method for Transformer in the embodiments of the present invention;

[0052] Figure 8 Schematic diagram of the method for judging error accumulation;

[0053] Figure 9 Schematic diagram of the fine-tuning process in the embodiments of the present invention.

[0054] Figure 10 Schematic diagram of the decoding method flow in the embodiments of the present invention. Detailed implementation manners

[0055] To further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments and drawings. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.

[0056] Term explanation:

[0057] Transformer: A deep learning network based on the attention mechanism proposed by Google, first proposed in an article "Attention is All You Need" published in NeurIPS (Conference on Neural Information Processing Systems) in 2017, and is a very popular model architecture.

[0058] GOP (Group of Pictures): In the GOP of the generalized end-to-end encoding and decoding method, generally the first frame is an I frame, and the remaining frames are P frames. The I frame is a reference frame. When encoding and decoding the I frame, the information of other frames is not required, while when encoding and decoding the P frame, the information of the previous frame is required. The transmission cost of the P frame is relatively low, but it cannot be decoded if the previous frame is lost.

[0059] Context mining model: A model for mining the information of the previous frame required for encoding the current frame. There is often a context mining model in the P frame compression model of the generalized end-to-end video encoding and decoding method.

[0060] Entropy encoding model: A model for predicting the distribution of latent features, and the prediction result will be used for entropy encoding. The quality of the prediction will affect the size of the video after encoding, and this model will be used at the encoding end and the decoding end.

[0061] Such as Figure 3As shown in the figure, a video encoding and decoding method combining video neural representation and generalization end-to-end encoding and decoding proposed in this embodiment includes an encoding method and a decoding method. The encoding method is implemented at the encoding end. The encoding method optimizes the encoding end model using the full-parameter fine-tuning method and optimizes the decoding end model (the decoding end model is the model to be used by the decoding method) using the efficient fine-tuning method, and converts the video into a latent feature bitstream and an efficient fine-tuning model parameter bitstream. The decoding method is implemented at the decoding end, receives the latent feature bitstream and the efficient fine-tuning model parameter bitstream, restores the bitstream into latent features and an efficient fine-tuning model, and decodes the video using the decoding end model and the efficient fine-tuning model.

[0062] Now assume that there is a generalization end-to-end codec trained on a large-scale dataset, that is, there are already a pre-trained encoding end model and a decoding end model.

[0063] Embodiment 1

[0064] As Figure 4 shown, the implementation of this embodiment of the encoding method includes the following processes:

[0065] S1: According to the network structure of the generalization end-to-end video encoding and decoding method, add and create an efficient fine-tuning model in its decoding end model.

[0066] This step mainly determines the structure of the model, that is, how to add the efficient fine-tuning model to the decoding end of the generalization end-to-end video codec. From the model architecture of the generalization end-to-end video codec, the model architectures of existing methods can be divided into models based on convolutional networks and models based on Transformers. Different architectures require different efficient fine-tuning methods. In this embodiment, a technology will be used as an example respectively, and other similar efficient fine-tuning technologies can also be applied here. The process of this step is as Figure 5 shown:

[0067] S11: Determine whether there is a decoding end model based on a convolutional network. If so, create an efficient fine-tuning model such as Conv-Adapter (Conv-Adapter: Exploring Parameter Efficient Transfer Learning for ConvNets, an efficient fine-tuning model specifically designed for convolutional networks), and connect the efficient fine-tuning model to the decoding end model based on the convolutional network. If not, skip this step.

[0068] The Conv-Adapter mainly consists of two convolutional layers. The first convolutional layer is responsible for reducing the feature dimension, and its parameters are initialized with a Gaussian distribution. Reducing the feature dimension can reduce the number of parameters in the convolutional network, and thus reduce the number of parameters in the efficient fine-tuning model. The second convolutional layer is responsible for changing the feature dimension to C, and its parameters are initialized with zeros. The value of C depends on which module in the decoding-end model the Conv-Adapter optimizes, and it is necessary to ensure that C is the same as the output feature dimension of this module. Initializing with zeros is to make the output of the efficient fine-tuning model be 0 at the beginning of the optimization, and temporarily have no impact on the decoding-end model, so that the optimization can start from a better starting point.

[0069] The following will combine Figure 6 to introduce the connection method between the Conv-Adapter and the decoding-end model based on the convolutional network:

[0070] There are two ways to introduce the Conv-Adapter. One is the parallel way shown by A and B in Figure 6 and the other is the serial way shown by Figure 6 C. As shown in Figure 6 A, when fine-tuning a certain convolutional layer in the parallel way, the input and output feature dimensions of the Conv-Adapter remain the same as those of this layer, and the output of the Conv-Adapter will be added to the output of this layer. As shown in Figure 6 B, when fine-tuning multiple convolutional layers in the parallel way, the input feature dimension of the Conv-Adapter remains the same as that of the first layer, the output feature dimension is the same as that of the last layer, and the output of the Conv-Adapter will be added to the output of the last layer. As shown in Figure 6 C, when fine-tuning in the serial way, the Conv-Adapter receives the output of a certain convolutional layer, then calculates the result with the same feature dimension, and adds this result to the output of this convolutional layer. In different encoding and decoding methods, which specific layers to fine-tune can be determined through some comparative fine-tuning experiments, subject to the actual effect. In addition, different efficient fine-tuning models can also be tried.

[0071] For example, for a generalization end-to-end codec based on the convolutional network DCVC-HEM (Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression, a hybrid spatio-temporal entropy coding model for video compression), the decoding-end model of DCVC-HEM includes a context mining model, a decoder model, and an entropy coding model. These models are all based on the convolutional network. Therefore, in this embodiment, as shown in Figure 6In this way, add Conv-Adapter to these models. It should be noted that there are many variants of convolutional networks. In the decoding-end models of DCVC-HEM, many models are not composed of a series of convolutional layers stacked serially, but are composed of a series of convolutional modules stacked together. Each module may contain several convolutional layers and is connected together in ways such as residual connections. For the sake of convenience, in this embodiment, a module is regarded as an ordinary convolutional layer for fine-tuning, and all use parallel fine-tuning, that is Figure 6 A or B in

[0072] S12: Determine whether there is a decoding-end model based on Transformer. If so, create an efficient fine-tuning model such as VPT (Visual Prompt Tuning), create an efficient fine-tuning model for the fully connected layer, and connect the efficient fine-tuning model to the decoding-end model based on Transformer. If not, skip this step.

[0073] VPT can achieve efficient fine-tuning of the Transformer-based network, such as Figure 7 shown Figure 7 What is shown in Figure 7 is the basic structure of the Transformer model, which consists of many Transformer layers and the final Head layer. The input of each layer is a series of features. The original input of the model is marked in black, that is Figure 7 x and e in

[0074] For the output layer after Transformer (that is, the fully connected layer that analyzes the Transformer output and generates the final result of the model), this embodiment proposes an efficient fine-tuning model similar to Conv-Adapter for fine-tuning the fully connected layer. Specifically, replace the convolutional layer in Conv-Adapter with a fully connected layer. The first fully connected layer is responsible for reducing the channel dimension, and then the second fully connected layer changes the feature dimension to a specified value C. C also depends on the connection method with the fully connected layer, which is exactly the same as S11 specifically.

[0075] Note that the effects of adding VPT features to different Transformer layers are different, and it is necessary to determine through control experiments which layers are most effective in adding prompts. There are also other efficient fine-tuning methods for Transformer architecture models, such as LoRA, Adapter Tuning (Parameter-Efficient Transfer Learning for NLP), etc.

[0076] Taking VCT (VCT: A Video Compression Transformer, a Transformer for video compression) as an example of the basic generalization end-to-end video codec. The entropy coding model of VCT adopts the Transformer architecture. For the entropy coding model, in this embodiment, in addition to the original input (i.e., latent features), randomly initialized optimizable VPT features are introduced. After these features are optimized, they can be regarded as a kind of prompt, providing information about the video to be encoded. This embodiment maintains the same features for all frames, that is, the same prompt is used for encoding and decoding all frames.

[0077] S2: Use the video to be encoded as the input of the optimization process, and optimize the encoding end model of the efficient fine-tuning model and the generalization end-to-end video codec method.

[0078] The specific process of this step is as follows:

[0079] S21: Design the objective function used in the optimization model according to whether there is error accumulation.

[0080] The objective function is the goal of optimizing the model and is an evaluation of the performance of the model in encoding and decoding videos. The process of optimizing the model is a process of continuously updating the neural network parameters according to the feedback of the objective function. Therefore, the design of the objective function will play a decisive role in the optimization effect. In this embodiment, different objective functions will be designed according to whether there is error accumulation in the generalization codec method. The specific process is as follows.

[0081] 1) Judge whether there is error accumulation according to the model connection method of the generalization end-to-end codec method.

[0082] Figure 8The method for judging error accumulation is shown as follows. The judgment method is as follows: when the generalization end-to-end video codec encodes each frame, each model needs to perform a calculation. If there is a model whose output when encoding frame t affects its input when encoding frame t+1, a loop link will be formed, resulting in error accumulation. For example, the output of the decoder model of DCVC-HEM based on the convolutional network (i.e., the decoded video) will be input into the context mining model to generate context features, and the context features will be input into the decoder model when encoding the next frame. In this way, the errors in the decoded video will be propagated to the decoder model through the context mining model, and then reflected in the video decoded from the next frame. This process will continue to cycle, and the errors will accumulate continuously starting from the I frame. However, all models of VCT are not affected by the output of the model itself, so there is no error accumulation.

[0083] 2) Design the corresponding objective function according to whether there is error accumulation.

[0084] If there is error accumulation, it means that the more frames are encoded, the more serious the error accumulation will be. Therefore, the objective function should consider the performance of the model when encoding multiple frames, so as to better measure the true performance of the model in actual encoding and decoding. In this embodiment, when considering encoding and decoding M frames at the same time, the corresponding objective function is:

[0085]

[0086] Among them, is the reconstruction loss of the decoded video, is the average number of bits used to transmit each pixel (BPP, bits per pixel), λ is used to balance the quality of the video frame and the number of bits to be transmitted, and is set according to actual needs. μ i is used to weight the reconstruction loss of the i-th frame with a weight of 1 + μ·i (μ>0). In this way, the weight of the reconstruction loss when encoding the next frame can be increased, that is, the objective function pays more attention to the quality of the next frame, so as to guide the model to appropriately improve the quality of the next frame and offset some of the accumulated errors.

[0087] When actually encoding a video, the generalized end-to-end encoding and decoding method generally divides a video into several GOPs. Error accumulation only exists within each GOP and will not accumulate to the next GOP. Therefore, the closer the number of M is to the size of the GOP, the better. However, the larger M is, the more memory space of the computing device is required. Due to device limitations, M is often much smaller than the actual GOP size. To better consider error accumulation, in this embodiment, a sliding window is used to intercept consecutive M frames of video starting from the first frame of the GOP, calculate the objective function once for these M frames, then update the model parameters once, and then the window is shifted one position as a whole until all frames within the GOP have been used for optimization. For example, if a GOP contains 3 frames and a window contains 2 frames (i.e., M = 2), then two calculations will be performed, and the frames intercepted twice are the first and second frames, and the second and third frames respectively.

[0088] Without error accumulation, the minimum number of frames required when sampling and encoding 1 frame in each optimization iteration, and the rate-distortion loss of a single frame can be directly used as the objective function:

[0089]

[0090] The meanings of each symbol are kept consistent with the objective function with error accumulation.

[0091] S22: Optimize the encoding-end model and the efficient fine-tuning model according to the designed objective function.

[0092] Figure 9 The following is a schematic diagram of the optimization principle of the encoding and decoding model. The specific process is as follows:

[0093] S221: Calculate the value of the objective function by simulating the encoding and decoding process.

[0094] There are mainly two parts to be calculated in the objective function. One part is the reconstruction distortion, and the other part is the average number of bits required to transmit each pixel. These can all be calculated by simulating the encoding and decoding process, which can also be called the forward propagation process.

[0095] The calculation process of the reconstruction distortion is as Figure 9 shown. The original video is converted into latent features by the encoder model. After the latent features are pseudo-quantized, they are then jointly decoded into the decoded video by the decoder model and the efficient fine-tuning model. Finally, the mean square error function or the SSIM (Mean Structural Similarity Index Measure) value between the decoded video and the original video is the reconstruction distortion.

[0096] The average number of bits required to transmit each pixel is obtained by the entropy coding model analyzing and calculating the pseudo - quantized latent features. The specific process is as follows: the entropy coding model estimates the distribution probability of the pseudo - quantized latent features, then estimates the number of bits required for entropy coding the latent features based on the probability, and finally divides the number of bits by the number of pixels.

[0097] S222: After obtaining the objective function value, use the gradient descent algorithm to optimize the encoding - end model and the efficient fine - tuning model at the decoding - end.

[0098] This process is the process of backpropagating the gradient of the objective function. The direction of backpropagation is exactly opposite to the direction of calculating the objective function. The gradients propagated to each model are used to update the model parameters. In this process, in this embodiment, all parameters of the encoding - end models such as the encoder model, the encoding - end part of the context mining model, and the encoding - end part of the entropy coding model are optimized, while the efficient fine - tuning module undergoes quantization - aware training, that is, the model optimization process. Quantization - aware training means pseudo - quantizing the parameters of the efficient fine - tuning model, and then using the pseudo - quantized parameters as the actual parameters to participate in the optimization process. This is because the parameters of the efficient fine - tuning model need to be quantized before transmission, and quantization will lead to a decrease in model performance. If quantization is performed during model optimization and the model is allowed to learn to use the quantized parameters, this problem can be better solved.

[0099] The specific method of quantization - aware training is as follows: First, normalize the model parameters to (-1, 1) through the tanh() function. Then, uniformly quantize the non - zero initialization layers of the efficient fine - tuning model and non - uniformly quantize the zero initialization layers through the STE (Straight - through estimator). Use the quantized model parameters as the parameters participating in the optimization.

[0100] It should be noted that both S221 and S222 describe the situation of calculating the objective function for a single frame. If the objective function involves multiple frames, the encoding - decoding processes of multiple frames need to be simulated, and both forward and backward propagations are also performed. Here, no repetitive description is made.

[0101] S223: Repeat the above two steps until the encoding - decoding performance of the model tends to be stable and the optimization is completed.

[0102] To determine whether the encoding - decoding performance of the model tends to be stable, it can be to check whether the relative change rate of the objective function is lower than a certain threshold. If it is lower than a certain threshold, it means that the model is basically converged and the optimization ends. It can also be to limit the optimization time. When the time ends, directly end the optimization, and use the model parameters with the optimal objective function as the finally fine - tuned model parameters.

[0103] S3: The encoding end obtains latent features through the optimized model, quantizes and entropy-encodes the latent features into a bitstream.

[0104] First, the encoding end converts each frame of the video into latent features through the optimized encoder model in S2. If the end-to-end encoding and decoding method based on the embodiment requires the use of context information, the encoder model will also use the output of the optimized context mining model in S2 when converting video frames.

[0105] Then, the encoding end quantizes the latent features and entropy-encodes the quantized latent features according to the distribution probability estimated by the optimized entropy encoding model in S22 to obtain the bitstream of the latent features, and transmits it to the decoding end.

[0106] It should be noted that some corresponding latent features of the entropy encoding model and the context mining model may also need to be quantized and entropy-encoded into bitstreams, and these bitstreams are transmitted to the user, such as DCVC-HEM, while VCT does not. For more details of this process, such as how to specifically use each model, reference can be made to the end-to-end encoding and decoding method based on this embodiment.

[0107] S4: The encoding end quantizes and entropy-encodes the parameters of the efficient fine-tuning model into a bitstream.

[0108] This embodiment introduces an efficient fine-tuning model to achieve fine-tuning of the decoding end model. The parameters of the efficient fine-tuning model are first obtained at the encoding end in S2, and then these parameters need to be transmitted from the encoding end to the decoding end to play their due role. Therefore, compared with the existing methods, this embodiment needs to additionally quantize and entropy-encode these parameters. The quantization method is consistent with the quantization-aware training mentioned in S22, and the entropy encoding method can adopt traditional entropy encoding techniques, such as Huffman coding and arithmetic coding. There are two reasons why this embodiment does not train a dedicated entropy encoding model to estimate the distribution probability of the parameters of the efficient fine-tuning model. One is that the entropy encoding model requires a large amount of parameters of the efficient fine-tuning model as data to be trained, which requires a very high training cost. The other is that the number of parameters of the efficient fine-tuning model is not much compared with the latent features and can basically be ignored, and the expected benefit is not high. Therefore, there is no need to spend too much cost to improve the entropy encoding performance.

[0109] Embodiment 2

[0110] This embodiment of the decoding method is also based on a generalized end-to-end encoding and decoding method trained on a large-scale dataset, that is, there is already a pre-trained decoding end model. As Figure 10 shown, it includes the following processes:

[0111] P1: The decoding end receives the bitstream of the latent features and the bitstream of the parameters of the efficient fine-tuning model, and restores the bitstreams to the efficient fine-tuning model and the latent features.

[0112] P11: After the decoding end receives all the high-efficiency fine-tuning parameter bitstreams, it restores the high-efficiency fine-tuning model and combines the high-efficiency fine-tuning model with the decoding-end model of the generalized end-to-end encoding and decoding method.

[0113] After the decoding end receives the high-efficiency fine-tuning parameter bitstream, according to the specified entropy coding method such as Huffman coding, arithmetic coding, etc., it restores the bitstream to the parameters of the high-efficiency fine-tuning model. These parameters may include the features of VPT, the parameters of Conv-Adapter, etc. If it includes a high-efficiency fine-tuning model that needs to be combined with the decoding end, such as Conv-Adapter, then according to the agreed model structure, it loads the high-efficiency fine-tuning model parameters into the model and connects them to the decoding-end model at the specified position in the way agreed by the encoding end and the decoding end. Such as the parallel connection method of Conv-Adapter.

[0114] P12: Use the entropy coding model of the generalized end-to-end encoding and decoding method that combines the high-efficiency fine-tuning model to restore the latent feature bitstream to latent features.

[0115] Use the entropy coding model that combines the high-efficiency fine-tuning model to restore the latent feature bitstream to latent features according to the latent feature restoration method of the generalized end-to-end encoding and decoding method. These latent features may include the latent features of video frames, the latent features of the context mining model, the latent features of entropy coding, which depends on the generalized end-to-end encoding and decoding method based on this embodiment.

[0116] P2: The decoding end uses the decoding model of the generalized end-to-end encoding and decoding method that combines the high-efficiency fine-tuning model to decode the latent features into video.

[0117] After restoration, use the decoding-end model that combines the high-efficiency fine-tuning model to decode the latent features into video. As described in P12, the latent features used here may include the latent features of video frames, the latent features of the context mining model, etc. The specific decoding process directly follows the generalized end-to-end encoding and decoding method based on this embodiment, but note that the model used should be replaced with the decoding-end model that combines the high-efficiency fine-tuning model.

[0118] For more details of the decoding method embodiment, such as how to specifically use each model, reference can be made to the end-to-end encoding and decoding method based on the embodiment.

[0119] Based on the same inventive concept, the embodiments of the present invention also provide a video encoding device and a video decoding device that combine video neural representation and generalized end-to-end video encoding and decoding.

[0120] Embodiment 3

[0121] The video encoding device that combines video neural representation and generalized end-to-end encoding and decoding in this embodiment includes:

[0122] A processor; a memory for storing the video to be encoded and neural network model parameters;

[0123] A signal transmitter;

[0124] And several programs running on the processor and the signal transmitter:

[0125] 1. Model optimization program: The processor uses this program to read the video to be encoded and optimize the encoding - end model and the efficient fine - tuning model of the generalized end - to - end coding and decoding method with this video.

[0126] 1) Generalized end - to - end coding and decoding model loading sub - program: The processor uses this program to create a preset generalized end - to - end coding and decoding model and load the corresponding model parameters in the memory.

[0127] 2) Efficient fine - tuning model creation and connection sub - program: The processor uses this program to create a preset efficient fine - tuning model, initialize the efficient fine - tuning model, and connect the efficient fine - tuning model to the decoding - end model of the generalized end - to - end coding and decoding model.

[0128] 3) Coding and decoding simulation sub - program: The processor uses this program to load the video to be encoded from the memory, sample video images, calculate the value of a set target function through the forward propagation process (the simulation process of encoding the video), calculate the gradient based on this value, and optimize the parameters of the encoding end of the end - to - end coding and decoding model and the parameters of the efficient fine - tuning model according to the backpropagation and gradient descent algorithms. This program runs in a loop until the change rate of the target function reaches the threshold set by the program.

[0129] 2. Encoding program: The processor uses this program to quantize and entropy - code the latent features and the parameters of the efficient fine - tuning model into a bitstream and send it to the signal sending device.

[0130] 1) Latent feature quantization and entropy - coding sub - program: The processor uses the program specified in the generalized end - to - end coding and decoding method to perform quantization and entropy - coding of the latent features.

[0131] 2) Efficient fine - tuning parameter quantization and entropy - coding sub - program: The processor quantizes the parameters of the efficient fine - tuning model, uses uniform quantization for the parameters initialized with Gaussian, uses non - uniform quantization for the parameters initialized with zero, and then uses an entropy - coding algorithm, such as Huffman coding, to encode the parameters into a bitstream.

[0132] 3. Signal sending program: The signal transmitter uses this program to send the bitstream.

[0133] Embodiment 4

[0134] The video decoding device combining video neural representation and generalized end - to - end coding and decoding in this embodiment includes:

[0135] A processor;

[0136] A memory for storing decoded video and neural network model parameters;

[0137] A signal receiver;

[0138] And several programs running on the processor and the signal receiver:

[0139] 1. Signal receiving program: The signal receiver receives a signal and transmits the code stream to the processor.

[0140] 2. Decoding program: The processor uses this program to restore the code stream to an efficient fine-tuning model and latent features. The processor uses the decoding-end model and the efficient fine-tuning model to decode the latent features into video.

[0141] 1) Generalized end-to-end codec model loading subroutine: The decoding-end processor creates a preset efficient fine-tuning model and a generalized end-to-end codec model, and loads the corresponding generalized end-to-end codec model parameters in the memory.

[0142] 2) Efficient fine-tuning parameter restoration subroutine: The decoding-end processor uses this program to restore the efficient fine-tuning parameters according to the entropy coding algorithm specified by the generalized end-to-end codec method, and loads the parameters into the efficient fine-tuning model.

[0143] 3) Latent feature restoration subroutine: The decoding-end processor uses this program to restore the latent features by using the entropy coding model created in the generalized end-to-end codec model loading subroutine combined with the efficient fine-tuning model.

[0144] 4) Latent feature inverse transformation subroutine: The decoding-end processor uses this program to restore the latent features to decoded video by using the decoding-end model created in the generalized end-to-end codec model loading subroutine combined with the efficient fine-tuning model.

[0145] 5) Storage program: The decoding-end processor uses this program to store the decoded video in the memory to complete decoding.

[0146] For the specific implementation manners of the components or programs of the above device, reference can be made to the corresponding steps of the video codec method combining the foregoing video neural representation and generalized end-to-end video codec, which will not be elaborated herein.

[0147] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A video coding method combining video neural representation and end-to-end encoding and decoding of generalization, characterized in that Including: 1) Based on the network structure of the generalized end-to-end video coding and decoding method, an efficient fine-tuning model is added and created in its decoding end model; 2) The video to be encoded is used as the input of the optimization process to optimize the encoding end model of the efficient fine-tuning model and the generalized end-to-end video coding and decoding method; 3) The encoding end obtains latent features through the optimized model, and quantizes and entropy-codes the latent features into a bitstream; 4) The encoding end quantizes and entropy-codes the parameters of the efficient fine-tuning model into a bitstream.

2. The video encoding method according to claim 1, wherein The optimization of the encoding end model and the efficient fine-tuning model further includes: 1) Design the objective function used in fine-tuning according to whether there is error accumulation; 2) Optimize the encoding end model and the efficient fine-tuning model according to the designed objective function.

3. The video encoding method according to claim 2, wherein The optimization of the encoding end model and the efficient fine-tuning model according to the designed objective function further includes: 1) Calculate the value of the objective function by simulating the coding and decoding process; 2) After obtaining the value of the objective function, use the gradient descent algorithm to optimize the encoding end model and the efficient fine-tuning model at the decoding end; 3) Repeat the above two steps until the coding and decoding performance of the model tends to be stable and the optimization is completed.

4. A video decoding method combining video neural representation and end-to-end encoding and decoding of generalization, characterized in that Including: 1) The decoding end receives the latent feature bitstream and the bitstream of the parameters of the efficient fine-tuning model, and restores the bitstream to the efficient fine-tuning model and the latent features; 2) The decoding end uses the decoding end model of the generalized end-to-end coding and decoding method combined with the efficient fine-tuning model to decode the latent features into a video.

5. The video decoding method according to claim 4, wherein The restoration of the bitstream to the efficient fine-tuning model and the latent features further includes: 1) After the decoding end receives all the bitstreams of the efficient fine-tuning parameters, it restores the efficient fine-tuning model parameter model, and combines the efficient fine-tuning model and the decoding end model of the generalized end-to-end coding and decoding method; 2) Use the entropy coding model of the generalized end-to-end coding and decoding method combined with the efficient fine-tuning model to restore the latent feature bitstream to the latent features.

6. A video coding device that combines video neural representation and end-to-end encoding and decoding for generalization, characterized in that Including: A processor; A memory for storing the video to be encoded and the neural network model parameters; A signal transmitter; And several programs running on the processor and the signal transmitter: Model optimization program: The processor uses this program to read the video to be encoded and optimize the encoding end model of the generalized end-to-end coding and decoding method and the efficient fine-tuning model with this video; Encoding program: The processor uses this program to quantize and entropy-code the latent features and the parameters of the efficient fine-tuning model into a bitstream and send it to the signal sending device; Signal sending program: The signal transmitter uses this program to send the bitstream to the receiving device.

7. The video encoding device according to claim 6, characterized in that, The model optimization program further includes: Generalized end-to-end coding and decoding model loading subprogram: The processor uses this program to create a pre-set generalized end-to-end coding and decoding model and load the corresponding model parameters in the memory; Efficient fine-tuning model creation and connection subprogram: The processor uses this program to create a pre-set efficient fine-tuning model, initialize the efficient fine-tuning model, and connect the efficient fine-tuning model to the generalized end-to-end coding and decoding model; Coding and decoding simulation subroutine: The processor uses this program to load the video to be encoded from the memory, sample video images, calculate the value of the set target function through the forward propagation process, calculate the gradient based on this value, and optimize the encoding end model parameters of the end-to-end coding and decoding method and the parameters of the efficient fine-tuning model according to the backpropagation and gradient descent algorithms.

8. The video encoding device according to claim 6, wherein The encoding program includes Latent feature quantization and entropy coding subroutine: The processor performs quantization and entropy coding of latent features using the program specified in the generalized end-to-end coding and decoding method; Efficient fine-tuning parameter quantization and entropy coding subroutine: The processor quantizes the parameters of the efficient fine-tuning model, and then encodes the parameters into a bitstream using the entropy coding algorithm.

9. A video coding device combining video neural representation and end-to-end encoding and decoding of generalization, characterized in that It includes: A processor; A memory for storing decoded videos and neural network model parameters; A signal receiver; And several programs running on the processor and the signal receiver: Signal reception program: The receiver receives signals and passes the bitstream to the processor; Decoding program: The processor uses this program to restore the bitstream to the efficient fine-tuning model and latent features. The processor uses the decoding end model and the efficient fine-tuning model to decode the latent features into a video.

10. The video decoding device according to claim 9, wherein, The decoding program includes: Generalized end-to-end coding and decoding model loading subroutine: The processor creates a preset efficient fine-tuning model and a generalized end-to-end coding and decoding model, and loads the corresponding generalized end-to-end coding and decoding model parameters in the memory; Efficient fine-tuning parameter restoration subroutine: The processor uses this program to restore the efficient fine-tuning parameters according to the entropy coding algorithm specified by the end-to-end coding and decoding method, and loads the parameters into the efficient fine-tuning model; Latent feature restoration subroutine: The processor uses this program to restore the latent features using the entropy coding model created in the generalized end-to-end coding and decoding model loading subroutine combined with the efficient fine-tuning model; Latent feature inverse transformation subroutine: The processor uses this program to restore the latent features to the decoded video using the decoding end model created in the generalized end-to-end coding and decoding model loading subroutine combined with the efficient fine-tuning model; Storage program: The processor uses this program to store the decoded video in the memory to complete the decoding.