Generative end-to-end video transmission

By using the UNet-Transformer diffusion model in video compression technology to generate some frames on the decoding end, the problem of high bit rate of video transmission in the prior art is solved, and video transmission with ultra-low bit rate is realized, while ensuring the consistency of video quality.

CN120151533APending Publication Date: 2025-06-13BEIJING UNION UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510087692.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing video compression technology needs to transmit all frames of information when transmitting video, resulting in high bit rate and making it difficult to achieve ultra-low bit rate video transmission.

Method used

The UNet-Transformer diffusion model is used to generate some frames at the decoding end, and through conditional guidance and space-time adaptive normalization mechanisms, the visual consistency between the generated frame and the original video is improved, thereby reducing the bit rate.

Benefits of technology

The bit rate of video transmission is significantly reduced without damaging video quality, and the limitations of traditional motion vector and residual transmission methods are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151533A_ABST
    Figure CN120151533A_ABST
Patent Text Reader

Abstract

A video decoding method employing a UNet neural network and a Transform neural network, comprising: parsing a first syntax element from a bitstream, the first syntax element indicating whether to enable neural network inter-frame prediction; in response to the first syntax element indicating that neural network inter prediction is enabled, a second syntax element is parsed from the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), where: inter prediction is performed on the N frames using the UNet DM (Diffusion Model), performing inter-frame prediction on a frame after the N frames by using a combination of UNet DM (Data Management) and Transform DM (Data Management); and in response to the first syntax element indicating that neural network inter prediction is enabled, parsing a third syntax element from the bitstream, the third syntax element indicating one of: (a) inter prediction is performed on frames following the N frames using UNet DM and Transform DM, respectively, to generate two frames, respectively, and selecting an optimal prediction frame based on a quality comparison between the two frames; or (b) respectively carrying out inter-frame prediction on the frames after the N frames by using UNet DM and Transform DM so as to respectively generate two frames, and carrying out frame fusion on the two frames so as to obtain a final prediction frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image and video processing, and more particularly, to a method, apparatus, and computer program product for generative end-to-end video transmission using UNet DM and Transformer DM. Background Art

[0002] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called "smart phones", video teleconferencing devices, video streaming devices, and the like.

[0003] Digital video devices implement video coding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) standards, ITU-T H.265 / High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC) (H.266), and extensions of such standards. By implementing such video coding techniques, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004] In April 2010, two major international video coding standard organizations, VCEG and MPEG, established the Joint Collaborative Team on Video Coding (JCT-VC) to jointly develop the High Efficiency Video Coding standard.

[0005] In 2013, JCT-VC completed the development of the High Efficiency Video Coding (HEVC) standard (also known as H.265), and subsequently released multiple versions.

[0006] To develop new technologies beyond HEVC, a new organization, the Joint Video Exploration Team, was established in 2015 and renamed the Joint Video Experts Team (JVET) in 2018. Based on HEVC, the research on Versatile Video Coding (VVC, H.266) was proposed by the JVET organization at the San Diego conference in the United States on April 10, 2018. It is a new generation of video coding technology improved on the basis of H.265 / HEVC. Its main goal is to improve the existing HEVC, provide higher compression performance, and optimize for emerging applications (360° panoramic video and high-dynamic range (HDR) video). The first version of VVC was completed in August 2020 and officially released as the H.266 standard on the ITU-T website.

[0007] Relevant documents and test platforms for HEVC and VVC can be obtained from https: / / jvet.hhi.fraunhofer.de / , and relevant proposals for VVC can be obtained from http: / / phenix.it-sudparis.eu / jvet / .

[0008] From HEVC, the application of neural networks to the video coding and decoding process has begun to be considered.

[0009] Among many neural networks, the Transformer model was initially proposed in Google's paper "Attention Is All You Need", which mainly involves: self-attention mechanism, encoder-decoder structure, residual connection and normalization. The Transformer model was initially designed for the field of natural language processing, but has also been widely used in the field of computer vision in recent years. By combining the ideas of convolutional neural networks (CNNs) and Transformer, more efficient tasks such as image classification and object detection can be achieved. This network structure of Transformer is only based on attention mechanisms and completely abandons the structures of loops and convolutions.

[0010] The self-attention mechanism is the core idea of Transformer, which allows the model to pay attention to information at different positions when processing sequence data. Specifically, the self-attention mechanism calculates the correlation between each position in the sequence and other positions to obtain an attention weight distribution, thereby achieving attention to information at different positions.

[0011] The Transformer adopts an encoder-decoder structure, where the encoder is responsible for converting the input sequence into a series of vector representations, and the decoder generates the output sequence based on these vector representations. This structure enables the Transformer to handle variable-length sequence data and has better generalization ability.

[0012] To address the problems of vanishing gradients and exploding gradients in deep networks, the Transformer introduces residual connections and normalization techniques. Residual connections allow the network to directly learn the residual function, thus alleviating the problem of vanishing gradients; while normalization makes the network more stable and easier to train by normalizing the data.

[0013] In the Transformer model, both the encoder and the decoder are composed of multiple stacked layers. Each layer consists of two sub-layers: the multi-head self-attention layer and the feed-forward neural network layer. The self-attention layer allows the model to weight and consider information at different positions when processing the input sequence, rather than relying solely on the sequential order of the sequence. It calculates attention weights to interact each position of the input sequence with other positions. Such an attention mechanism can capture important context information in the sequence and thus performs excellently in dealing with long-range dependencies. The feed-forward neural network layer independently maps and transforms the features at each position. It uses a fully connected feed-forward neural network and performs a non-linear transformation on the features through an activation function (such as ReLU). In the encoder, the input sequence is processed through multiple encoder layers, and each layer generates a new feature representation. The output of the encoder can be used for various downstream tasks, such as text classification, named entity recognition, etc.

[0014] The Unet neural network model is a convolutional neural network (CNN) widely used in medical image segmentation. It was initially proposed by Olaf Ronneberger et al. in the 2015 paper "U-Net: Convolutional Networks for Biomedical Image Segmentation". The Unet neural network model was initially proposed for image segmentation, and the flexibility of the Unet model makes it the first choice for many researchers and engineers and is thus applicable to the field of video coding and decoding.

[0015] The core of the Unet model is a symmetric "U" - shaped structure, which consists of a contracting path (encoder) and an expanding path (decoder). The contracting path is used to capture context information, and the expanding path is used to precisely locate the segmentation boundaries, which are connected by skip connections in the middle.

[0016] The encoder part consists of multiple convolutional blocks. Each convolutional block contains two 3x3 convolutional layers and a ReLU activation function, followed by a 2x2 max - pooling layer for downsampling. After each downsampling, the size of the feature map is halved and the depth is doubled.

[0017] Decoder: The decoder part also consists of multiple convolutional blocks, but here an upsampling operation is used to restore the image size. Upsampling is usually achieved through transposed convolution operations, doubling the size of the feature map. After each upsampling layer, the upsampled feature map is merged with the feature map of the corresponding encoder layer through concatenation, so that more detailed information can be retained.

[0018] The key innovation of UNet is the introduction of skip connections in the decoder, that is, connecting the feature maps in the encoder with the corresponding feature maps in the decoder. These skip connections can help the decoder better utilize feature information at different levels, thereby improving the accuracy of image segmentation and the ability to retain details.

[0019] Diffusion Models (DM) are models used for artificial intelligence generation. They gradually add Gaussian noise to turn it into pure Gaussian noise zzz, and then generate new images by gradually denoising zzz. Diffusion models are actually a specific application model of artificial neural networks, mainly used to generate various high - resolution images.

[0020] Diffusion models were proposed in 2015, and their motivation comes from non - equilibrium thermodynamics. By simulating a random diffusion process, they gradually transform random noise into the target data distribution, thereby generating new data samples.

[0021] Simply put, diffusion models are divided into two processes: "adding noise" and "denoising" (also known as the forward process and the reverse process).

[0022] Adding - noise process: Continuously add noise to the input data until it becomes pure Gaussian noise. At each moment, a part of Gaussian noise is added to the image. The latter moment is obtained by adding more noise to the previous moment. After T times of adding - noise operations, the input x 0 will continuously mix in Gaussian noise, and finally the image x T will become a pure - noise image that conforms to the standard normal distribution.

[0023] Denoising process: Starting from pure Gaussian noise, gradually remove the noise to obtain an image that satisfies the training data distribution. During the denoising process, we hope to train a neural network that can learn T denoising operations to transform x T back to x 0 . The learning objective of the network is to make the T denoising operations exactly cancel out the corresponding noise addition operations. After training, just randomly sample a noise from the standard normal distribution and then use the neural network in the reverse process to restore the noise into an image, and an image can be generated.

[0024] Diffusion models are particularly suitable for image generation or video generation. An image generation network will learn how to map a vector into an image. When designing the network architecture, the most important thing is to design the learning objective so that the images generated by the network are similar to the images in the given dataset. The variational autoencoder (VAE) approach is to use two networks, one that learns to encode an image into a vector and the other that learns to decode the vector back into an image, and their objective is to make the restored image as similar as possible to the original image. After learning, the decoder is the image generation network. Diffusion models are a more specific type of VAE. It fixes the encoding process as adding noise and lets the decoder learn how to remove each step of the previously added noise.

[0025] In Unet, through the layer-by-layer processing of the convolutional layers, the noise information will be gradually weakened during the feature extraction process, and the decoder can use the feature maps transmitted by the skip connection structure to restore the image details while avoiding information loss and blurring. Therefore, Unet is a conditional denoising network and can be used as a conditional denoising network for diffusion models. Thus, the Unet diffusion model has become a frequently used diffusion model for the image generation process.

[0026] The Transformer DM was initially proposed in 2023 (W. Peebles and S. Xie, “Scalable Diffusion Models with Transformers,” 2023 IEEE / CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 4172 - 4182, doi: 10.1109 / ICCV51070.2023.00387), known as Diffusion Transformers (DiTs), aiming to break through the bottlenecks of traditional methods. The input of DiT is a set of multi-channel latent variable feature maps (Noised Latent), usually with a size of 32×32×4 for example. These latent variables are generated by predicting the noise from the previous time step. Before entering the Transformer model, the latent variables need to be processed by the Patchify module. The core of the Patchify module is to divide the image or features into small patches, similar to dividing a painting into small jigsaw puzzles. The size of each patch is determined by the parameter p, such as 8×8 or 16×16. Subsequently, these patches are transformed into a one-dimensional Token sequence through a linear transformation (Embed) for further processing by the Transformer. In addition, the time step information (Timestep ttt) and class label (Class Label ccc) are used as conditional information, which are transformed into vector Tokens through an embedding operation and directly concatenated into the input Token sequence to provide additional context information for the model. In this way, Patchify realizes an efficient conversion from the image space to the sequence space, laying the foundation for subsequent Transformer modeling.

[0027] In recent years, video compression technology has been widely studied, and traditional video coding frameworks have also been continuously evolving. The new generation of ECM-based video coding technology further optimizes and improves each module. For example, Zhao et al. proposed a motion optimization method based on template matching, which significantly improved the performance of motion compensation, effectively increased the video compression efficiency, and can be applied to the ECM framework. X. Xie designed a set of optimized interpolation filters, which improved the accuracy of chrominance motion compensation, reduced the computational complexity at the same time, and enhanced the video coding efficiency. The CPGA network proposed by Qiang Zhu et al. enhanced the detail performance and reduced compression artifacts by aggregating coding prior guidance information, thus improving the visual quality of the compressed video.

[0028] With the development of deep learning, researchers have begun to explore how to replace some modules in traditional video coding frameworks with deep learning methods. J. Jia et al. proposed a deep learning method for generating reference frames, enhancing the inter-frame prediction in the VVC standard. By generating high-quality reference frames, the prediction accuracy was improved, which is an important innovation in the field of video coding. M.-J. Chen developed a hierarchical B-frame coding technique, which utilized adaptive feature modulation and deep learning methods to enhance the compression efficiency and quality of YUV 4:2:0 videos. Y. Mao proposed a neural network-based rate control method that can dynamically adjust coding parameters to improve the efficiency of VVC coding while optimizing video quality and bitrate.

[0029] In addition, researchers have also started to conduct end-to-end designs for video coding frameworks. The research by Y. Zhao integrated neural networks into traditional video coding, surpassing the Enhanced Compression Model (ECM). Optimization was carried out in multiple aspects such as prediction, transformation, and quantization, and the compression efficiency and video quality were improved through end-to-end optimization. J. Chen proposed a rate control strategy from sparse to dense, which optimized the balance between video quality and bandwidth by dynamically adjusting bitrate allocation. J. Zhou proposed an end-to-end distributed video coding method, which was optimized by combining distributed resources and neural networks, enhancing the compression efficiency and the robustness of network transmission.

[0030] In the past few years, traditional residual-based video compression technologies (such as MPEG-2, AVC / H.264, HEVC / H.265, and VVC / H.266) have developed to a mature stage. At the same time, end-to-end video compression methods based on conditional coding have emerged one after another, achieving significant bitrate reduction. However, with the continuous growth of video data demand, the potential for further reducing the bitrate is gradually limited. The progress of generative technologies in the field of image and video processing has brought new possibilities and challenges to video compression. Generative technologies achieve high-quality image and video generation through conditional or unconditional guidance. In recent years, video generation frameworks have introduced various guidance methods, including text, texture maps, and single-frame guidance (such as SVD). Although these models have achieved significant improvements in video quality metrics, their generated results still have diversity and randomness, limiting the precision and controllability of generation.

[0031] In addition, whether it is an end-to-end design or a traditional coding framework, it usually requires transmitting all the information of the original video, which leads to a high bitrate requirement. The framework we designed innovatively generates some frames at the decoding end, which means that not all frames need to be transmitted, thus significantly reducing the transmission bitrate. Summary of the Invention

[0032] Due to the booming development of generative models, we adopted a strategy of hybrid generative models and designed a cascaded diffusion model based on the VAE generative model. Its backbone consists of a cascade of Unet and Transformer. The frames generated by this model have a high similarity to the original video at the encoding end, and at the same time, the generated images are selected and optimized. This framework is applicable to any type of video transmission, can efficiently process long and short videos, and performs well in video transmissions with frequent scene mutations and transitions.

[0033] The present invention explores a novel video compression method combining generative techniques: at the decoding end, the UNet-Transformer diffusion model (DM) is utilized to generate subsequent frames based on the guidance of partial frames, thereby further reducing the bit rate.

[0034] More specifically, the present invention proposes a conditional guidance diffusion model that integrates Unet for local feature capture and Transformer for long-range dependency modeling, and enhances the visual consistency between the generated frames and the original video frames through a spatio-temporal adaptive normalization mechanism. To solve numerous technical problems in the prior art, a super-low bit rate video generation framework (KGVT) of the UNet-Transformer diffusion model (DM) is proposed, overcoming the limitations of traditional motion vector (MV) and residual transmission methods. In the encoding stage, first, some original frames are compressed according to specific requirements, and the number of compressed frames is determined. Then, the UNet-Transformer DM model is used to generate subsequent frames, and the generated frames are evaluated by a quality comparator to select the frame closest to the reference frame. At the same time, a threshold discriminator generates a transmission vector to decide whether to continue transmitting the original frame. In the decoding stage, the received transmission vector and the decoded frame jointly guide the UNet-Transformer DM model to generate corresponding subsequent frames, and the generated frames are integrated to reconstruct the complete video sequence.

[0035] In addition, the present invention designs a semantic mutation frame extraction network to improve the generation accuracy. We observe that some frames in the video exhibit significant semantic differences. For this reason, we design a semantic mutation frame extraction network specifically for identifying and transmitting frames with significant semantic changes. This framework not only addresses the challenges of super-low bit rate compression but also ensures semantic consistency between adjacent frames. A brief mind map of the proposed framework is as Figure 1 shown.

[0036] According to one aspect, a video decoding method using a UNet neural network and a Transformer neural network includes:

[0037] Parsing a first syntax element from a bitstream, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0038] In response to the first syntax element indicating enabling neural network inter-frame prediction, a second syntax element is parsed from the bitstream, where the second syntax element indicates N frames after an intra-coded frame (Intra frame), and:

[0039] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model),

[0040] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0041] In response to the first syntax element indicating enabling neural network inter-frame prediction, a third syntax element is parsed from the bitstream, where the third syntax element indicates one of the following:

[0042] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on the quality comparison between the two frames; or

[0043] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain the final predicted frame.

[0044] In a preferred aspect, in the Transformer DM, the concatenation of conditional features and noise latent features is used as the input, where a conditional spatio-temporal mask is used to unify the conditional input (I):

[0045] I = F(1 - M) + CM

[0046] where C represents the conditional video frame, F represents the noise, and M represents the spatio-temporal mask.

[0047] In a preferred aspect, the method further includes: parsing the parameters of the UNet DM and the Transformer DM from the bitstream.

[0048] In another aspect, a video encoding method using a UNet neural network and a Transformer neural network includes:

[0049] Set a first syntax element in the bitstream, where the first syntax element indicates whether to enable neural network inter-frame prediction;

[0050] When the first syntax element is set to indicate enabling neural network inter-frame prediction, a second syntax element is set in the bitstream, and the second syntax element indicates N frames after an intra-coded frame (Intra frame), where, in the decoding loop:

[0051] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model);

[0052] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0053] When the first syntax element is set to indicate enabling neural network inter-frame prediction, a third syntax element is set in the bitstream, and the third syntax element indicates one of the following:

[0054] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on the quality comparison between the two frames; or

[0055] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final predicted frame.

[0056] In a preferred aspect, in Transformer DM, the concatenation of conditional features and noise latent features is used as the input, where a conditional spatio-temporal mask is used to unify the conditional input (I):

[0057] I = F(1 - M) + CM

[0058] where C represents the conditional video frame, F represents the noise, and M represents the spatio-temporal mask.

[0059] In a preferred aspect, the method further includes: setting the parameters of the UNet DM and the Transformer DM in the bitstream.

[0060] In a preferred aspect, the method further includes: using optical flow to determine a semantic mutation frame as the next Intra frame after the Intra frame.

[0061] In yet another aspect, a video decoding device employing a UNet neural network and a Transformer neural network includes:

[0062] One or more memories configured to process frame data to be decoded and decoded frame data; and

[0063] One or more processing units, configured to decode frame data to be decoded, and further configured to:

[0064] Parse a first syntax element from the bitstream, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0065] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a second syntax element from the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), where:

[0066] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model);

[0067] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0068] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a third syntax

[0069] element from the bitstream, the third syntax element indicating one of the following:

[0070] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on a quality comparison between the two frames; or

[0071] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final predicted frame.

[0072] In yet another aspect, a video decoding chip employing a UNet neural network and a Transformer neural network includes:

[0073] An input terminal configured to receive frame data to be decoded, the frame data to be decoded being at least a part of an encoded bit rate;

[0074] An output terminal configured to output the decoded frame data for display;

[0075] One or more memories configured to process the frame data to be decoded and the decoded frame data; and

[0076] One or more processing units configured to decode the frame data to be decoded, and further configured to:

[0077] Parse a first syntax element from the frame data to be decoded, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0078] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a second syntax element from the frame data to be decoded, the second syntax element indicating N frames after an intra-coded frame (Intra frame), where:

[0080] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model),

[0081] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0082] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, decode from the frame data to be decoded

[0083] Parse a third syntax element, the third syntax element indicating one of the following:

[0084] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on the quality comparison between the two frames; or

[0085] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final predicted frame.

[0086] In another aspect, a computer program product includes a non-transitory storage medium storing code for performing the method according to the present invention. Description of the Drawings

[0087] Figure 1 Shows an embodiment of a semantic mutation frame extraction network.

[0088] Figure 2 Shows the overall framework of the UNet-Transformer DM model according to the present invention.

[0089] Figure 3 Shows a flowchart of the video decoding method according to the present invention.

[0090] Figure 4 Shows a flowchart of the video encoding method according to the present invention.

[0091] Figure 5A block diagram of a video decoding device according to an embodiment of the present invention is shown.

[0092] Figure 6 A block diagram of a video decoding chip according to an embodiment of the present invention is shown.

[0093] Figure 7 A skeletal diffusion model of UNet and Transformer according to an embodiment of the present invention is shown. Detailed implementation manners

[0094] Now, various solutions will be described with reference to the accompanying drawings. In the following description, for the purpose of explanation, a number of specific details are set forth in order to provide a thorough understanding of one or more solutions. However, it is obvious that these solutions can be implemented without these specific details.

[0095] As used in this application, the terms "component", "module", "system", etc. are intended to refer to computer-related entities, such as but not limited to, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be but is not limited to: a process running on a processor, a processor, an object, an executable, an execution thread, a program, and / or a computer. For example, an application running on a computing device and the computing device can both be components. One or more components can be located within an execution process and / or an execution thread, and a component can be located on one computer and / or distributed on two or more computers. Additionally, these components can execute from various computer-readable media having various data structures stored thereon. Components can communicate with each other by means of local and / or remote processes, such as according to a signal having one or more data packets, for example, data from a component that interacts with another component in a local system, a distributed system, and / or interacts with other systems over a network such as the Internet by means of a signal.

[0096] Note that the terms "encoder" and "decoder" are used in the present invention. In a video coding and decoding framework, "encoder" and "decoder" refer to devices for compressing and decompressing video; while in the field of artificial neural networks, "encoder" and "decoder" refer to the "encoder" and "decoder" in a neural network, rather than the devices for compressing and decompressing video in a video coding and decoding framework.

[0097] Figure 2 An overall framework of the UNet-Transformer DM model according to the present invention is shown.

[0098] First, the functions of the video generation framework (LGVF) of the present invention will be introduced.

[0099] At the encoding end, first, the decoding end is simulated to perform k-frame decoding, and UNet-Trans DM is used to predict the subsequent t frames from k to k+t. However, as the number of prediction steps increases, the quality of the predicted frames may decline. At this time, the threshold discriminator will judge these t frames to decide whether to transmit the original frames and generate the corresponding label vectors, thus forming a cyclic feedback. Considering the semantic mutations of video frames, we extract the mutation frames of the video to prevent the unpredictability of sudden events. At the decoding end, after receiving the label vectors and key frames, these two parts jointly guide the generation process of the entire video. In this process, we make full use of the generation mechanism to achieve low bitrate compression and visual consistency.

[0100] In the generation stage, the first T frames are predicted by UNet-DM, while the subsequent frames are predicted by UNet-DM and Trans-DM simultaneously. The reason for this processing method is that Transformer can capture long-range dependencies and flexibly allocate attention to different parts of the input sequence, thus more effectively capturing the inter-frame information. Therefore, in the case of a large number of conditional frames, rich temporal and spatial information, Trans-DM shows better performance. Finally, the quality comparator evaluates these two groups of prediction results, selects the frames with the best quality, and passes these frames into the threshold discriminator for the final discrimination.

[0101] Secondly, the processing of key frames is introduced. Here, the key frame refers to the frame that needs to be intra-coded.

[0102] In video processing, due to the occurrence of drastic transitions or sudden events in the video, the optical flow value will increase significantly. Such semantic mutation frames are usually difficult to accurately predict by the prediction model. To solve this problem, this paper designs an optical flow-based semantic mutation frame extraction network.

[0103] Specifically, the optical flow fraction can be calculated by the magnitude of the optical flow vector. For the optical flow vector of each pixel

[0104]

[0105] where u and v represent the velocities on the x-axis and y-axis respectively, that is, the optical flow components.

[0106] The optical flow fraction of the entire frame is calculated by averaging the optical flow magnitudes of all pixels or using other aggregation methods:

[0107]

[0108] where N is the total number of pixels in the current frame, u i and v iis the optical flow component of the i-th pixel. When the optical flow fraction of a certain frame exceeds the preset threshold, it is extracted, as shown in the schematic diagram. In the subsequent generation stage, if there is a mutation frame at the corresponding moment, the semantic mutation frame is used for replacement, so as to ensure that the generated video is highly consistent with the original video visually. Through this method of extracting semantic mutation frames based on optical flow, sudden inter-frame changes in the video can be effectively dealt with, and the effects of video compression and generation can be improved.

[0109] Subsequently, we introduce the usage of UNet-Transformer DM in the present invention.

[0110] In the generation stage, we designed a skeleton diffusion model that combines UNet and transformer, which has two modes.

[0111] Figure 7 Shows the skeleton diffusion model of UNet and Transformer according to an embodiment of the present invention.

[0112] As Figure 7 Shown in the lower part, the first mode is the cascade mode. In this mode, the model uses the UNet-based backbone network for prediction in the first few frames, and then UNet and transformer jointly perform prediction in the subsequent frames. Among the generated frames, the quality of the two is compared, and the best frame is selected for replacement.

[0113] As Figure 7 Shown in the upper part, the second mode is the parallel mode. In this mode, UNet and transformer simultaneously predict the subsequent frames, then perform frame fusion on the predicted frames, and finally splice the fused frames to form a complete video sequence.

[0114] The experimental results show that the first cascade mode performs more superiorly in terms of video generation quality. This indicates that in the video generation process, adopting the strategy of single model prediction in the first few frames plus multi-model quality comparison in the subsequent frames can more effectively improve the visual consistency and generation quality of the video.

[0115] Then, we introduce the quality comparator used when using the skeleton diffusion model of UNet and Transformer.

[0116] When two diffusion models generate 5 frames at the same time, we select the optimal frame by comparing the quality metrics of the frames at each corresponding moment. Since a single image quality metric cannot comprehensively evaluate the image quality, we propose a criterion for selecting the optimal frame, as shown in the following formula:

[0117]

[0118] This standard selects the frames with higher p-values by calculating and comparing the p-values of each frame, combines them into the corresponding t frames at that moment, and uses them as the output of the quality comparator.

[0119] Subsequently, we introduce some details of the diffusion model.

[0120] The initially proposed diffusion model processes images through a noise addition and denoising process, which follows a Markov chain. Therefore, the model cannot predict future states. However, after introducing the prediction function of UNet, the model has the ability of conditional prediction. Even so, there is still room for improvement in the model's ability to generate videos with visual consistency. For example, video generation guided by text or abstract images usually exhibits a large degree of diversity and uncontrollability. In video compression and transmission, the past, present, and future frames of the video are all known. Under these known conditions, the decoding end should restore a specific video with the least amount of data transmission possible.

[0121] Assume x 0 is an instance in the data sample. We infer the distribution of the sample x t-1 at the current moment through the conditional distribution of the sample x t at the previous moment, and thus obtain the forward process formula of the diffusion model during the generation process:

[0122]

[0123] where β t is the variance, then is the mean, where α t = 1 - β t , that is, α t is the square of the mean.

[0124] The generation of new samples can be achieved by starting from the Gaussian noise xT and solving the reverse diffusion process (RDP).

[0125]

[0126] When modeling, we constructed the following diffusion model formula with the previous few frames as conditions:

[0127]

[0128] where P represents the sequence of past frames, and X0 is the sequence of the current frame. By using the past frames to predict the current frame, this method can perform video generation and compression transmission more precisely.

[0129] This UNet-DM adopts a model structure with UNet as the backbone (as Figure 1As shown, its core function lies in the spatio-temporal adaptive normalization technology used in the normalization layer. By dynamically adjusting past and future frames, this method can better ensure the transmission of temporal dynamics between different frame blocks. This means that our network can learn an implicit model of spatio-temporal dynamics, thereby providing more accurate information for frame generation.

[0130] This technology enables the model to effectively capture the spatio-temporal dependencies between frames in the video, enhancing the quality and consistency of video generation. At the same time, spatio-temporal adaptive normalization ensures that the details of the generated video are closer to the original video by adjusting the information distribution in each frame block, thus improving the model's generation ability.

[0131] In the transformer diffusion model, we introduce a spatio-temporal masking mechanism. The input data can be latent features of pure noise or a concatenation of conditional features and noise latent features. Next, we use a conditional spatio-temporal mask to unify the conditional input I, as shown in the following equation:

[0132] I = F(1 - M) + CM

[0133] Where C represents the conditional video frame, F represents the noise, and M represents the spatio-temporal mask.

[0134] As shown in the figure, when inputting latent features, these features are first encoded in terms of time and position. The encoded features are integrated into x, thereby transmitting time and position information to the model. A transformer with a cascaded structure is adopted in the model, with a total of 28 layers, and each layer contains a temporal attention module, a spatial attention module, and a multi-head attention mechanism. Finally, the denoised predicted frame is reconstructed in the last layer. Experimental results show that the model can effectively generate and predict target video frames.

[0135] Through these designs, our model has successfully captured spatio-temporal information during video generation, significantly improving the generation quality.

[0136] Figure 3 The flowchart of the video decoding method according to the present invention is shown. The video decoding method can be implemented by a dedicated video codec or a general-purpose processor.

[0137] In step 301, the decoder receives the encoded bitstream to be decoded. In one embodiment, the bitstream can be received from the network via a network port. In another embodiment, the bitstream can be received from an external memory via various input interfaces, and these external memories can include, for example, hard disk drives, network disks, solid-state memories, optical disk memories, magnetic tapes, another computing device, and so on.

[0138] The decoder can be various devices, apparatuses, chips, etc. for decoding video data. In this decoder, a video decoding method using a UNet neural network and a Transformer neural network according to various embodiments of the present invention can be implemented.

[0139] In step 303, the method may include: parsing a first syntax element from the bitstream, where the first syntax element indicates whether neural network inter-frame prediction is enabled. For example, the first syntax element may be set to 0 to indicate that neural network inter-frame prediction is not enabled; or it may be set to 1 to indicate that neural network inter-frame prediction is enabled.

[0140] In step 305, the method may include: parsing a second syntax element from the bitstream.

[0141] In one embodiment, the method may include: in response to the first syntax element indicating that neural network inter-frame prediction is enabled, parsing a second syntax element from the bitstream. The second syntax element indicates N frames after an intra-coded frame (Intra frame). The decoder may use the UNet DM (diffusion model) to perform inter-frame prediction on the N frames, and use a combination of UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames.

[0142] In one embodiment, the first syntax element and the second syntax element are different syntax elements. The decoder may, in response to the first syntax element indicating that neural network inter-frame prediction is enabled (e.g., having a value of "0"), parse the second syntax element from the bitstream.

[0143] In another embodiment, the first syntax element and the second syntax element may be the same syntax element. For example, when the value of this syntax element is 0, it indicates that neural network inter-frame prediction is not enabled; while when the value of this syntax element is a non-zero value, it may indicate that neural network inter-frame prediction is enabled, and the specific value of this syntax element serves as the value of the second syntax element, indicating N frames after an intra-coded frame (Intra frame).

[0144] In step 307, the method may include: in response to the first syntax element indicating that neural network inter-frame prediction is enabled, parsing a third syntax element from the bitstream, where the third syntax element indicates one of the following:

[0145] (a) Using UNet DM and Transformer DM respectively to perform inter-frame prediction on the frames after the N frames to generate two frames respectively, and selecting the best predicted frame based on the quality comparison between the two frames; or

[0146] (b) Use UNet DM and Transformer DM respectively to perform inter-frame prediction on the frames after the N frames, to generate two frames respectively, and perform frame fusion on the two frames to obtain a final predicted frame.

[0147] In an optional preferred step 309, the method may include: parsing the parameters of the UNet DM and the Transformer DM from the bitstream.

[0148] In one aspect, in a preferred embodiment of step 307, in Transformer DM, the concatenation of conditional features and noise latent features is used as the input, wherein a conditional spatio-temporal mask is used to unify the conditional input (I):

[0149] I = F(1 - M) + CM

[0150] Wherein, C represents a conditional video frame, F represents noise, and M represents a spatio-temporal mask.

[0151] Figure 4 The flowchart of the video coding method according to the present invention is shown. The video decoding method can be implemented by a dedicated video codec or a general-purpose processor.

[0152] In step 401, the method may include: receiving a video frame to be encoded. For example, raw video data frames captured in real time can be received from a digital camera via various wired or wireless interfaces. For another example, pre-stored video data frames to be encoded can be received from various external or internal storage devices capable of storing video data via various wired or wireless interfaces. These video data frames to be encoded can be raw video data frames previously captured by a digital camera, or can be computer-generated computer graphics video data frames, such as various forms of animations. The above-mentioned external or internal memory may include, for example, a hard disk drive, a network disk, a solid-state memory, an optical disk memory, a magnetic tape, another computing device, and the like.

[0153] In step 403, the method may include: setting a first syntax element in the bitstream, and the first syntax element indicates whether neural network inter-frame prediction is enabled.

[0154] In step 405, the method may include: setting a second syntax element in the bitstream.

[0155] In one embodiment, the method may include: when the first syntax element is set to indicate enabling neural network inter-frame prediction, setting a second syntax element in the bitstream. The second syntax element indicates N frames after an intra-coded frame (Intra frame). In the decoding loop, the UNet DM (diffusion model) may be used to perform inter-frame prediction on the N frames, and a combination of UNet DM and Transformer DM may be used to perform inter-frame prediction on the frames after the N frames.

[0156] In one embodiment, the first syntax element and the second syntax element are different syntax elements. The encoder may set the value of the second syntax element in the bitstream when the first syntax element is set to indicate enabling neural network inter-frame prediction (e.g., having a value of "0").

[0157] In another embodiment, the first syntax element and the second syntax element may be the same syntax element. For example, when the value of the syntax element is 0, it indicates not enabling neural network inter-frame prediction; while when the value of the syntax element is a non-zero value, it may indicate enabling neural network inter-frame prediction, and the specific value of the syntax element serves as the value of the second syntax element, indicating N frames after an intra-coded frame (Intra frame).

[0158] In step 407, the method may include: when the first syntax element is set to indicate enabling neural network inter-frame prediction, setting a third syntax element in the bitstream, where the third syntax element indicates one of the following:

[0159] (a) Using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames respectively to generate two frames, and selecting the best predicted frame based on the quality comparison between the two frames; or

[0160] (b) Using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames respectively to generate two frames, and performing frame fusion on the two frames to obtain a final predicted frame.

[0161] In an optional preferred step 409, the method may include: setting the parameters of the UNet DM and the Transformer DM in the bitstream.

[0162] In one aspect, in a preferred embodiment of step 307, in Transformer DM, the concatenation of conditional features and noise latent features is used as the input, where a conditional spatio-temporal mask is used to unify the conditional input (I):

[0163] I = F(1 - M) + CM

[0164] Among them, C represents a conditional video frame, F represents noise, and M represents a spatio-temporal mask.

[0165] In one aspect, the method may further include ( Figure 4 not shown in the figure): using optical flow to determine a semantic mutation frame as the next Intra frame after the Intra frame.

[0166] Figure 5 FIG. shows a block diagram of a video decoding device 500 according to an embodiment of the present invention. In one embodiment, the video decoding device may be a dedicated video decoding device, for example, a video decoding chip, a video decoding circuit, or may be a configured video decoding device, such as a field programmable gate array (FPGA), a general-purpose video decoding processing chip, a GPU, etc. In another embodiment, the video decoding device may be a video codec device, including various dedicated video codecs (CODECs), general-purpose processing devices.

[0167] As Figure 5 shown, the video decoding device 500 includes: one or more memories 501 configured to process frame data to be decoded and decoded frame data; and one or more processor units 503 configured to decode the frame data to be decoded and further configured to perform operations corresponding to each step of the decoding method according to Figure 3 the description.

[0168] Figure 6 FIG. shows a block diagram of a video decoding chip 600 according to an embodiment of the present invention. In one embodiment, the video decoding chip may be a dedicated video decoding chip. In another embodiment, the video decoding chip may be a video codec (CODEC) chip.

[0169] As Figure 5 shown, the video decoding device 600 includes one or more memories 601 and one or more processor units 603. In one embodiment, the one or more memories 601 are cache memories (CACHE) provided inside the chip, so it may not be able to store data of more than one group of pictures (GOP), and may even only be able to store a part of the data of one frame.

[0170] As Figure 5 shown, the video decoding device 600 includes an input end 605 and an output end 607.

[0171] In one embodiment, the input end 605 and the output end 607 are a set of the same or different chip pins.

[0172] The input terminal 605 is configured to receive frame data to be decoded, and the frame data to be decoded is at least a part of the encoded bit rate. The output terminal 607 is configured to output the decoded frame data for display.

[0173] The one or more memories 601 are configured to process the frame data to be decoded and the decoded frame data.

[0174] The one or more processing units are configured to decode the frame data to be decoded and are further configured to perform operations corresponding to respective steps of the decoding method according to Figure 3 the decoding method described.

[0175] According to another aspect, the present disclosure may also relate to a computer program product for performing the methods described herein. According to a further aspect, the computer program product has a non-transitory storage medium on which computer code / instructions are stored, which when executed by a processor can implement the various operations described herein.

[0176] When implemented in hardware, the video encoder can be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but alternatively, the processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors and a DSP core, or any other such configuration. Additionally, at least one processor can include one or more modules operable to perform one or more of the above steps and / or operations.

[0177] When implementing the video encoder using hardware circuits such as ASICs, FPGAs, etc., it can include various circuit blocks configured to perform various functions. Those skilled in the art can design and implement these circuits in various ways according to various constraints imposed on the entire system to implement the various functions disclosed in the present invention.

[0178] The various embodiments herein are listed in the following numbered items:

[0179] Item 1, A video decoding method using a UNet neural network and a Transformer neural network, comprising:

[0180] Parsing a first syntax element from a bitstream, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0181] In response to the first syntax element indicating enabling neural network inter-frame prediction, a second syntax element is parsed from the bitstream, where the second syntax element indicates N frames after an intra-coded frame (Intra frame), where:

[0182] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model).

[0183] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0184] In response to the first syntax element indicating enabling neural network inter-frame prediction, a third syntax element is parsed from the bitstream, where the third syntax element indicates one of the following:

[0185] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on the quality comparison between the two frames; or

[0186] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final predicted frame.

[0187] Item 2. The method according to Item 1, wherein in the Transformer DM, the concatenation of conditional features and noise latent features is used as the input, and a conditional spatio-temporal mask is used to unify the conditional input (I):

[0188] I = F(1 - M) + CM

[0189] where C represents a conditional video frame, F represents noise, and M represents a spatio-temporal mask.

[0190] Item 3. The method according to any one of Items 1 - 2, further comprising:

[0191] Parse the parameters of the UNet DM and the Transformer DM from the bitstream.

[0192] Item 4. A video coding method using a UNet neural network and a Transformer neural network, comprising:

[0193] Set a first syntax element in the bitstream, where the first syntax element indicates whether to enable neural network inter-frame prediction;

[0194] When the first syntax element is set to indicate enabling neural network inter-frame prediction, a second syntax element is set in the bitstream, and the second syntax element indicates N frames after an intra-coded frame (Intra frame), where, in the decoding loop:

[0195] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model),

[0196] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0197] When the first syntax element is set to indicate enabling neural network inter-frame prediction, a third syntax element is set in the bitstream, and the third syntax element indicates one of the following:

[0198] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best prediction frame based on the quality comparison between the two frames; or

[0199] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final prediction frame.

[0200] Item 5. The method according to item 4, wherein, in Transformer DM, the concatenation of conditional features and noise latent features is used as the input, and a conditional spatio-temporal mask is used to unify the conditional input I:

[0201] I = F(1 - M) + CM

[0202] where C represents the conditional video frame, F represents the noise, and M represents the spatio-temporal mask.

[0203] Item 6. The method according to any one of items 4-5, further comprising:

[0204] Set the parameters of the UNet DM and the Transformer DM in the bitstream.

[0205] Item 7. The method according to any one of items 4-6, comprising:

[0206] Use optical flow to determine a semantic mutation frame as the next Intra frame after the Intra frame.

[0207] Item 8. A video decoding device adopting a UNet neural network and a Transformer neural network, comprising:

[0208] One or more memories configured to process frame data to be decoded and decoded frame data; and

[0209] One or more processing units configured to decode the frame data to be decoded and further configured to:

[0210] Parse a first syntax element from a bitstream, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0211] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a second syntax element from the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), where:

[0212] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model),

[0213] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0214] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a third syntax element from the bitstream, the third syntax element indicating one of the following:

[0215] The third syntax element indicates one of the following:

[0216] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best predicted frame based on a quality comparison between the two frames; or

[0217] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final predicted frame.

[0218] Item 9. A video decoding chip using a UNet neural network and a Transformer neural network, comprising:

[0219] An input end configured to receive frame data to be decoded, the frame data to be decoded being at least a part of an encoded bit rate;

[0220] An output end configured to output decoded frame data for display;

[0221] One or more memories configured to process the frame data to be decoded and the decoded frame data; and

[0222] One or more processing units configured to decode the frame data to be decoded and further configured to:

[0223] Parse a first syntax element from the frame data to be decoded, the first syntax element indicating whether neural network inter-frame prediction is enabled;

[0224] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, parse a second syntax element from the frame data to be decoded, the second syntax element indicating N frames after an intra-coded frame (Intra frame), where:

[0226] Perform inter-frame prediction on the N frames using the UNet DM (diffusion model),

[0227] Perform inter-frame prediction on the frames after the N frames using a combination of UNet DM and Transformer DM; and

[0228] In response to the first syntax element indicating that neural network inter-frame prediction is enabled, decode from the frame data to be decoded

[0229] Parse a third syntax element, the third syntax element indicating one of the following:

[0230] (a) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and select the best prediction frame based on the quality comparison between the two frames; or

[0231] (b) Perform inter-frame prediction on the frames after the N frames using UNet DM and Transformer DM respectively to generate two frames, and perform frame fusion on the two frames to obtain a final prediction frame.

[0232] Item 10. A computer program product, including a non-transitory storage medium storing code for performing the method according to any one of Items 1-7.

[0233] Although the foregoing disclosure documents discuss exemplary solutions and / or embodiments, it should be noted that many changes and modifications can be made herein without departing from the scope of the described solutions and / or embodiments defined by the claims. Moreover, although the elements of the described solutions and / or embodiments are described or claimed in the singular, plural cases can also be contemplated unless explicitly stated to be limited to the singular. Additionally, all or part of any solution and / or embodiment can be used in combination with all or part of any other solution and / or embodiment unless otherwise indicated.

Claims

1. A video decoding method using a UNet neural network and a Transformer neural network, comprising: Parsing a first syntax element from a bitstream, the first syntax element indicating whether neural network inter prediction is enabled; In response to the first syntax element indicating that neural network inter-frame prediction is enabled, a second syntax element is parsed from the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), wherein: Using the UNet DM (diffusion model) to perform inter-frame prediction on the N frames, Using a combination of UNet DM and Transformer DM to perform inter-frame prediction on frames after the N frames; and In response to the first syntax element indicating that neural network inter prediction is enabled, a third syntax element is parsed from the bitstream, the third syntax element indicating one of the following: (a) using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames respectively to generate two frames respectively, and selecting the best predicted frame based on quality comparison between the two frames; or (b) Using UNet DM and Transformer DM, respectively, inter-frame prediction is performed on the frames after the N frames to generate two frames respectively, and the two frames are frame-fused to obtain a final predicted frame.

2. The video decoding method according to claim 1, wherein: In Transformer DM, the concatenation of conditional features and noisy latent features is used as input, where the conditional input is unified using a conditional spatiotemporal mask (I): I=F(1-M)+CM Among them, C represents the conditional video frame, F represents the noise, and M represents the spatiotemporal mask.

3. The video decoding method according to claim 1 or 2, further comprising: Parameters of the UNet DM and the Transformer DM are parsed from the bitstream.

4. A video encoding method using a UNet neural network and a Transformer neural network, comprising: Setting a first syntax element in a bitstream, the first syntax element indicating whether neural network inter prediction is enabled; When the first syntax element is set to indicate that neural network inter prediction is enabled, a second syntax element is set in the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), wherein, in a decoding loop: Using the UNet DM (diffusion model) to perform inter-frame prediction on the N frames, Using a combination of UNet DM and Transformer DM to perform inter-frame prediction on frames after the N frames; and When the first syntax element is set to indicate that neural network inter prediction is enabled, a third syntax element is set in the bitstream, the third syntax element indicating one of the following: (a) using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames respectively to generate two frames respectively, and selecting the best predicted frame based on quality comparison between the two frames; or (b) Using UNet DM and Transformer DM, respectively, inter-frame prediction is performed on the frames after the N frames to generate two frames respectively, and the two frames are frame-fused to obtain a final predicted frame.

5. The video encoding method according to claim 4, wherein: In Transformer DM, the concatenation of conditional features and noisy latent features is used as input, where the conditional input I is unified using the conditional spatiotemporal mask: I=F(1-M)+CM Among them, C represents the conditional video frame, F represents the noise, and M represents the spatiotemporal mask.

6. The video encoding method according to claim 4 or 5, further comprising: Parameters of the UNet DM and the Transformer DM are set in the bitstream.

7. The video encoding method according to any one of claims 4 to 6, comprising: The optical flow is used to determine a semantic mutation frame as the next Intra frame after the Intra frame.

8. A video decoding device using a UNet neural network and a Transformer neural network, comprising: One or more memories configured to process frame data to be decoded and decoded frame data; as well as One or more processing units configured to decode frame data to be decoded, and further configured to: Parsing a first syntax element from a bitstream, the first syntax element indicating whether neural network inter prediction is enabled; In response to the first syntax element indicating that neural network inter-frame prediction is enabled, a second syntax element is parsed from the bitstream, the second syntax element indicating N frames after an intra-coded frame (Intra frame), wherein: Using the UNet DM (diffusion model) to perform inter-frame prediction on the N frames, Use a combination of UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames; as well as In response to the first syntax element indicating that neural network inter prediction is enabled, a third syntax element is parsed from the bitstream, the third syntax element indicating one of the following: (a) Use UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames, respectively. to generate two frames separately and select the best prediction frame based on the quality comparison between the two frames; or (b) using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames, Two frames are generated respectively, and the two frames are fused to obtain the final predicted frame.

9. A video decoding chip using a UNet neural network and a Transformer neural network, comprising: An input terminal configured to receive frame data to be decoded, the frame data to be decoded being at least a portion of the encoded bit rate; an output terminal configured to output the decoded frame data for display; One or more memories configured to process the frame data to be decoded and the decoded frame data; as well as One or more processing units are configured to decode the frame data to be decoded, and are further configured to: Parsing a first syntax element from the frame data to be decoded, the first syntax element indicating whether neural network inter-frame prediction is enabled; In response to the first syntax element indicating that neural network inter-frame prediction is enabled, a second syntax element is parsed from the frame data to be decoded, wherein the second syntax element indicates N frames after an intra-coded frame (Intra frame), wherein: Using the UNet DM (diffusion model) to perform inter-frame prediction on the N frames, Use a combination of UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames; as well as In response to the first syntax element indicating that neural network inter-frame prediction is enabled, a third syntax element is parsed from the frame data to be decoded, the third syntax element indicating one of the following: (a) Use UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames, respectively. to generate two frames separately and select the best prediction frame based on the quality comparison between the two frames; or (b) using UNet DM and Transformer DM to perform inter-frame prediction on the frames after the N frames, Two frames are generated respectively, and the two frames are fused to obtain the final predicted frame.

10. A computer program product, comprising a non-transitory storage medium, wherein the non-transitory storage medium stores codes for executing the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Transform-based high-power face video super-resolution processing method

    CN120765460A

  • High-magnification face video super-resolution processing method based on transformer

    CN120765460B