Extensible resolution neural network for generative video compression

By using generative video coding technology, deep generative models and advanced video coding techniques, the problems of compression redundancy and bandwidth imbalance in complex scenarios of existing video coding are solved, and efficient video reconstruction and bandwidth utilization are achieved.

CN121842387APending Publication Date: 2026-04-10ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video coding technologies suffer from limitations in feature design and generation schemes when dealing with complex scenes such as moving human bodies, leading to compression redundancy and uneven bandwidth requirements. Furthermore, traditional hybrid codecs perform poorly in complex scenarios.

Method used

Generative video coding (GVC) technology is adopted, which uses a deep generative model for video compression, reconstructs video frames through neural networks, adjusts the network width and depth to adapt to different resolution requirements, and combines advanced video coding techniques such as VVC to achieve flexible resolution output.

Benefits of technology

It improves the ultra-low bit rate and high-quality reconstruction capabilities of video communication, reduces compression redundancy, enhances the ability to handle complex scenes, and achieves more efficient bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842387A_ABST
    Figure CN121842387A_ABST
Patent Text Reader

Abstract

A video decoding method comprises: decoding an image bit stream associated with a video sequence to reconstruct a key frame of the video sequence and obtain an extraction feature of the reconstructed key frame; decoding a feature bit stream associated with the video sequence to obtain extracted features of one or more inter-frames of the video sequence; obtaining motion information and occlusion information based on the extracted features of the reconstructed key frame and the extracted features of the one or more inter-frame frames; resampling the reconstructed key frame by a neural network based on the motion information and the shielding information; and reconstructing the video sequence by the neural network based on the resampled reconstruction key frame, and adjusting the network width and the network depth of the neural network according to the input resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 705,037, filed October 9, 2024, entitled "Resolution-Expandable Generator for Generative Video Compression," which is incorporated herein by reference in its entirety. This application also claims priority to U.S. Patent Application No. 19 / 323,240, filed September 9, 2025. Technical Field

[0002] This disclosure relates generally to video processing, and more specifically, to a scalable resolution generator (e.g., a generative neural network) for generative video compression. Background Technology

[0003] Video consists of a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, while the decompression process is usually called decoding. There are many video coding formats that use standardized video coding techniques, the most common being based on prediction, transform, quantization, entropy coding, and loop filtering. Standardization organizations have developed video coding standards that specify particular video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Universal Video Coding (VVC / H.266) standard, and the AVS standard. As more and more advanced video coding technologies are incorporated into video standards, the coding efficiency of new video coding standards is also increasing. Summary of the Invention

[0004] According to some embodiments, a video decoding method includes: decoding an image bitstream associated with a video sequence to reconstruct keyframes of the video sequence and obtain extracted features of the reconstructed keyframes; decoding a feature bitstream associated with the video sequence to obtain extracted features of one or more inter-frame frames of the video sequence; obtaining motion information and occlusion information based on the extracted features of the reconstructed keyframes and the extracted features of the one or more inter-frame frames; resampling the reconstructed keyframes by a neural network based on the motion information and occlusion information; and reconstructing the video sequence by the neural network based on the resampled reconstructed keyframes, adjusting the network width and network depth of the neural network according to the input resolution.

[0005] According to some embodiments, a video coding method includes: encoding an image bitstream, the image bitstream including encoded information for keyframes of a video sequence, wherein the image bitstream can be decoded to reconstruct the keyframes; and encoding a feature bitstream, the feature bitstream including encoded information for extracted features of one or more inter-frames of the video sequence. Using the features of the reconstructed keyframes and the features of the one or more inter-frames encoded in the feature bitstream, dense motion information and occlusion information are generated for a neural network to resample the reconstructed keyframes. The neural network reconstructs the video sequence by adjusting the network width and network depth according to the input resolution.

[0006] According to some embodiments, a method for storing an image bitstream and a feature bitstream includes: generating an image bitstream and a feature bitstream based on a video sequence, wherein the image bitstream includes: encoded information for reconstructing keyframes of the video sequence and obtaining extracted features of the reconstructed keyframes, and the feature bitstream includes: encoded information for obtaining extracted features of one or more inter-frame frames of the video sequence; and storing the image bitstream and the feature bitstream in at least one non-transitory computer-readable medium. The video sequence is reconstructed by a neural network that resamples the reconstructed keyframes. The neural network adjusts its width and depth according to the input resolution. Attached Figure Description

[0007] Embodiments and aspects of this disclosure are illustrated in the following detailed description and accompanying drawings. Various features shown in the figures are not drawn to scale.

[0008] Figure 1 This is a schematic diagram illustrating an example encoding / decoding process of a generative video coding (GVC) algorithm according to some embodiments of the present disclosure.

[0009] Figure 2 A schematic diagram illustrates an example framework for video compression in a video encoding system according to some embodiments of the present disclosure.

[0010] Figure 3 This is a schematic diagram illustrating an example architecture of an end-to-end deep video compression (DVC) framework according to some embodiments of the present disclosure.

[0011] Figure 4 This is a block diagram of an example apparatus for encoding or decoding image data according to some embodiments of the present disclosure.

[0012] Figure 5 This is a schematic diagram illustrating the basic framework of a depth-based generative video compression scheme based on a first-order motion model (FOMM) according to some embodiments of the present disclosure.

[0013] Figure 6 This is a schematic diagram illustrating an encoder for a depth-based video generative compression scheme based on Compact Feature Temporal Evolution (CFTE) according to some embodiments of the present disclosure.

[0014] Figure 7 This is a schematic diagram illustrating a decoder for a depth-based video generative compression scheme based on Compact Feature Temporal Evolution (CFTE) according to some embodiments of the present disclosure.

[0015] Figure 8 This is a schematic diagram illustrating an example network structure 800 of a scalable resolution neural network according to some embodiments of the present disclosure.

[0016] Figure 9 This is a schematic diagram illustrating an example generative video coding system with a scalable resolution neural network according to some embodiments of the present disclosure.

[0017] Figure 10 This is a flowchart of an example method for decoding a bitstream according to some embodiments of the present disclosure.

[0018] Figure 11 This is a flowchart of an example method for encoding a bitstream according to some embodiments of the present disclosure. Detailed Implementation

[0019] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise stated, the same numerals in different figures represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the invention. Rather, these embodiments are merely examples of apparatuses and methods consistent with the relevant aspects of the invention described in the appended claims. Specific aspects of this disclosure are described in more detail below. In the event of any conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0020] In the era of AI-Generated Content (“AIGC”), video coding technology is rapidly evolving towards more intelligent, immersive, and interactive applications. One key technology is Generative Video Coding (GVC), which leverages the powerful inference capabilities of deep generative models in visual data compression, achieving superior rate-distortion (RD) performance compared to traditional hybrid codecs such as High Efficiency Video Coding (HEVC) and General Video Coding (VVC). For example, some generative video codecs evolved from deep image animation methods, which represent high-dimensional visual input signals in a compact form and employ powerful deep generative models to achieve high-quality signal reconstruction / animation. For instance, deep animation codecs use 2D keypoint representations for ultra-low bitrate video conferencing. Similarly, in speaking face video coding, 3D keypoints are used for free-viewpoint control, while feature matrices can represent facial temporal trajectories in a more compact manner.

[0021] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing a universal video coding standard (VVC / H.266). The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0022] To achieve this goal, since 2015, JVET has been continuously developing technologies that surpass HEVC using the Joint Exploratory Model (JEM) reference software. With coding techniques incorporated into JEM, JEM achieved significantly higher coding performance than HEVC. In October 2017, VCEG and MPEG issued a Joint Call for Proposals (CfP), officially launching the development of a next-generation video compression standard that surpasses HEVC. Responses to the CfP were evaluated at the JVET meeting in San Diego in April 2018, and formal development of the VVC standard began in April 2018.

[0023] Since April 2018, the VVC standard has progressed smoothly and continues to incorporate more coding technologies that provide better compression performance. VVC adopts the hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0024] Video is a set of still images (or “frames”) arranged chronologically to store visual information. Video capture devices (e.g., cameras) can be used to capture and store those images chronologically, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display) can be used to display these images chronologically. Furthermore, in some applications, video capture devices can transmit the captured video in real time to the video playback device (e.g., a computer with a display), for example, for video observation, conferencing, or live streaming.

[0025] To reduce the storage space and transmission bandwidth required by these applications, the video can be compressed before storage or transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or by dedicated hardware. The module used for compression is typically called an "encoder," while the module used for decompression is typically called a "decoder." The encoder and decoder can be collectively referred to as a "codec." The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder can include circuit systems such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process embedded in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, the codec can decompress the video according to a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec can be called a "transcoder".

[0026] Video coding identifies and retains useful information for image reconstruction while ignoring information unimportant to the reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such a coding process can be called "lossy." Otherwise, it can be called "lossless." Most coding processes are lossy, a trade-off between reducing required storage space and transmission bandwidth.

[0027] Useful information about an image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes can include variations in pixel position, brightness, or color, with positional changes being of primary interest. The positional changes of a set of pixels characterizing an object can reflect the movement of that object between the reference image and the current image.

[0028] An image encoded without reference to another image (i.e., whose reference image is itself) is called an "I-image". An image is called a "P-image" if some or all blocks in an image (e.g., blocks of various parts of a video image) are predicted using intra-frame prediction or inter-frame prediction with a reference image (e.g., one-way prediction). An image is called a "B-image" if at least one block in an image is predicted using two reference images (e.g., two-way prediction).

[0029] Figure 1 This is a schematic diagram illustrating an example encoding / decoding process of a generative video coding (GVC) algorithm according to some embodiments of the present disclosure. For example, the encoding / decoding process can be used for generative face video coding. Figure 1 As shown, a generative video coding system 100 can be configured to compress and reconstruct an input video sequence 110, which has a key reference frame 112 and one or more inter-frames 114 following the key reference frame 112. In some embodiments, the generative video coding system 100 may include an encoder 120 and a decoder 130, each including multiple interconnected components designed to efficiently process video data. In some cases, the generative video coding system 100 may simultaneously utilize VVC coding techniques and advanced generative models to achieve flexible resolution output.

[0030] The encoder 120 of the generative video coding system 100 can process input video frames to generate an image bitstream 132 associated with the key reference frame 112 and a feature bitstream 134 associated with the inter-frame frame 114. Then, the decoder 130 of the generative video coding system 100 can use the image bitstream 132 and the feature bitstream 134 to reconstruct the video sequence to obtain an output video 150, which includes the decoded key reference frame 152 and the reconstructed inter-frame frame 154. Figure 1 As shown, in the encoder 120, the key reference frame 112 of the video can be compressed using various image / video codecs such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Universal Video Coding (VVC). Figure 1 In one embodiment, the VVC encoder 122 is configured to encode the key reference frame 112 into the image bitstream 132. Subsequent inter-frame frames 114 can be characterized as compact transmission symbols and encoded into the feature bitstream 134 using the analysis model 124 and parameter encoding module 126 in the encoder 120.

[0031] At the decoder end, the decoder 130 can decode the image bitstream 132 using the VVC decoder 142 to obtain the decoded key reference frame 152. Furthermore, through the corresponding parameter decoding module 144, the decoder 130 can decode the feature bitstream 134 to obtain compact facial information. The decoded key reference frame 152 and the compact facial information are jointly fed into the synthesis model 146 to obtain the reconstructed inter-frame frame 154 for reconstructing the video. In this way, ultra-low bitrate and high-quality reconstruction of video communication can be achieved.

[0032] In some embodiments, Figure 1 The capabilities of generative coding may be limited by feature design and generation scheme. For example, the generative video codec primarily employs explicit feature representations with materialization, leading to unnecessary compression redundancy. Furthermore, such representations may lack expressiveness and generalizability, failing to handle complex scenes, such as moving human figures. For instance, explicit features including landmarks, keypoints, and segmentation maps are used for low-bandwidth video chat compression. In some embodiments, different feature representations may result in different bandwidth requirements.

[0033] The development of video compression standards such as AVC, HEVC, and VVC aims to achieve high compression performance. These standards employ block-based hybrid video coding frameworks to utilize spatial, temporal, and entropy redundancy in video.

[0034] Figure 2 A schematic diagram of an example framework 200 for video compression in a video encoding system according to some embodiments of the present disclosure is shown. Typically, the video compression encoder generates the bitstream based on the input current frame. The decoder then reconstructs the video frames based on the received bitstream. Figure 2 The framework 200 described in the text follows a prediction-transformation architecture.

[0035] The input video is processed block by block. Specifically, the input frame x t The video is divided into blocks of the same size (e.g., 8×8), such as square regions. The encoding process of the video compression algorithm at the encoder end will be discussed below.

[0036] The input frame x tProcessed by a block-based motion estimation module 210, which is configured to estimate the current frame x t Compared with the previously reconstructed frame The motion between them. Based on the input frame x t and the previously reconstructed frame The block-based motion estimation module 210 outputs the motion vector v corresponding to each block. t Then, the corresponding motion vector v is processed by the motion compensation module 220. t So as to base the motion vector v defined in the motion estimation module 210 t By using the previously reconstructed frame The corresponding pixels in the image are copied to the current frame to obtain the prediction frame. Therefore, the original frame x is obtained. t With the predicted frame The residual r between t ,for In some embodiments, the motion compensation prediction performed above is also referred to as "inter-frame prediction", "inter-image prediction", or "temporal prediction".

[0037] In generating the residual r t Then, the encoder can convert the residual r t The results are fed into the transformation stage 232 and quantization stage 234 of the transformation and quantization module 230 to generate quantization results. In some embodiments, a linear transformation (e.g., DCT) may be used prior to the quantization to achieve better compression performance. Different transformation algorithms may use different basic modes. Various transformation algorithms may be used in the transformation stage 232, for example, discrete cosine transform, discrete sine transform, etc. The transformation in the transformation stage 232 is reversible. That is, the encoder can recover the residual r by performing the inverse operation of the transformation (referred to as the "inverse transform") by the inverse transform module 240. t For video coding standards, the encoder and corresponding decoder can use the same transform algorithm (and thus the same base mode). Therefore, the encoder can simply record the transform coefficients, and the decoder can reconstruct the residual r based on these coefficients. t Without needing to receive the basic mode from the encoder.

[0038] The encoder can further compress the transform coefficients in quantization stage 234. During the transform process, different basis modes can characterize different change frequencies (e.g., brightness change frequency). Because the human eye is generally better at identifying low-frequency changes, the encoder can ignore information about high-frequency changes without causing a significant degradation in decoding quality. For example, in quantization stage 234, the encoder can generate quantization residual coefficients by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. Following this operation, some transform coefficients of the high-frequency fundamental mode can be converted to zero, while the transform coefficients of the low-frequency fundamental mode can be converted to smaller integers. The encoder can ignore the zero-valued quantization residual coefficients. Therefore, the transformation coefficients are further compressed. The quantization process is also reversible, wherein the quantization residual coefficients... The transform coefficients can be reconstructed in the inverse operation of the quantization (referred to as "inverse quantization").

[0039] Because the encoder ignores the remainder of division in the rounding operation, the quantization stage 234 may be lossy. Typically, the quantization stage 234 may constitute the primary source of information loss during the encoding process. The greater the information loss, the lower the quantization residual coefficient. The fewer bits may be needed. To obtain different levels of information loss, the encoder can use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0040] The encoder can transmit the motion vector v t and the quantized residual coefficient The data is fed to the binary encoding module 250 to generate a bitstream to complete the forward path. Through the binary encoding module 250, the encoder can use binary encoding techniques to process the motion vector v. t Quantization residual coefficient Encoding is performed, for example, using binary encoding techniques such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding (CABAC), or any other lossless or lossy compression algorithm. Therefore, the motion vector v can be encoded using entropy coding methods. t and the quantized residual coefficient The bits are encoded into bits and then sent to the decoder.

[0041] As described above, the quantization result can be used through the inverse transform module 240. The reconstructed residual is obtained through the inverse transformation. In the process described above, after quantization stage 234, the encoder can quantize the residual coefficients. The data is fed into the inverse quantization and inverse transformation stages of the inverse transformation module 240 to generate the reconstructed residual. During the inverse quantization stage, the encoder can process the quantization residual coefficients. Inverse quantization is performed to generate reconstructed transform coefficients. During the inverse transform stage, the encoder can generate the reconstructed residual based on the reconstructed transform coefficients. Then, the encoder can reconstruct the residual Added to the prediction frame To obtain the reconstructed frame to be used in the next iteration process. Right now, The reconstructed frame The motion estimation module 210 will be used for motion estimation in the (t+1)th frame.

[0042] In generating the reconstructed frame The encoder can then apply a loop filter to the reconstructed frame. This reduces or eliminates distortion (e.g., blocking) introduced by the inter-frame prediction. In some embodiments, the encoder may apply various loop filtering techniques during the loop filtering stage, such as deblocking, Sample Adaptive Shift (SAO), Adaptive Loop Filtering (ALF), etc. In SAO, after the deblocking filtering, a nonlinear amplitude mapping is introduced within the inter-frame prediction loop to reconstruct the original signal amplitude using a lookup table described by a small number of additional parameters determined through histogram analysis at the encoder end.

[0043] The loop-filtered reference image can be stored in the decoded frame buffer 260 for later use (e.g., as an inter-frame prediction reference frame for a future frame of the video sequence). The encoder can store one or more reference frames in the buffer 260 for inter-frame prediction. In some embodiments, the encoder can adjust the parameters of the loop filter (e.g., loop filter strength) and the motion vector v during the encoding stage. t Quantization residual coefficient Other information is also encoded. The encoder iteratively executes the process described above to encode each frame of the video sequence.

[0044] For the decoder, based on the bits provided by the binary encoding module 250 in the encoder, corresponding motion compensation, inverse transform, and frame reconstruction operations can be performed to obtain the reconstructed frame.

[0045] Next, we will describe the end-to-end deep video compression (DVC) technique. Figure 3 This is a schematic diagram illustrating an example architecture of an end-to-end depth video compression (DVC) framework 300 according to some embodiments of the present disclosure.

[0046] With the development of deep learning, many deep learning-based algorithms can be introduced to replace or enhance video coding tools. These algorithms include intra / inter-frame prediction, entropy coding, and in-loop filtering.

[0047] Figure 3 The video compression framework 300 shown employs an end-to-end video compression depth model, which can jointly optimize some or all components of video compression, such as motion estimation, motion compression, and residual compression. Specifically, motion information can be obtained and the current frame reconstructed using learning-based optical flow estimation. Then, two autoencoder-like neural networks are used to compress the corresponding motion and residual information. These modules can be jointly learned through a single loss function, whereby they cooperate by weighing the relationship between reducing the number of compressed bits and improving the quality of the decoded video. Figure 2 The video compression frame 200 shown is... Figure 3 There is a one-to-one correspondence between the end-to-end depth video compression frameworks 300 shown. The relationship and differences between these two frameworks 200 and 300 will be discussed below.

[0048] In the framework 300, the motion estimation and compression module 310 includes an optical flow network 312, a motion vector (MV) encoder network 314, a quantization module 316, and a motion vector (MV) decoder network 318. The optical flow network 312 may be a convolutional neural network (CNN) model configured to estimate the optical flow, which is considered to be motion information v. t The MV encoder network 314 and the MV decoder network 318 do not directly encode the original optical flow value, but are instead configured to compress and decode the optical flow value, respectively. The motion representation output by the MV encoder network 314 is m. t And the quantized motion representation output by the quantization module 316 is Then, based on the quantified motion representation... The corresponding reconstructed motion information is decoded using the MV decoder network 318.

[0049] For the motion compensation process, the motion compensation network 320 is designed based on the reconstructed motion information. To obtain the predicted frame

[0050] For the aforementioned transformation and quantization process Figure 2 The linear transformation in the video compression framework 200 is replaced by a highly nonlinear residual encoder-decoder network, and the residual r t The residual encoder network 332 nonlinearly maps to the representation y. t Then, output y. t The quantization characterization is obtained by quantization module 334. In some embodiments, quantization methods can be used to construct an end-to-end training scheme. Then, the quantized representation... The residual is fed into the residual decoder network 340 to obtain the reconstructed residual.

[0051] During the entropy encoding phase, in the testing phase, the quantized motion representation and the residual characterization The bit rate estimation network (BRF) encodes 350 bits into bits and sends them to the decoder. During the training phase, the CNN can be used to estimate the bit rate increment. and The probability distribution of each symbol in the text.

[0052] Figure 3 The frame reconstruction process in the frame 300 and the buffer 360 used in the frame reconstruction process are described in the original text. Figure 2 The frame reconstruction process in the frame 200 is the same as that in the buffer 260, so for the sake of brevity, it will not be described again here.

[0053] Figure 4 This is a block diagram of an example apparatus 400 for encoding or decoding image data according to some embodiments of the present disclosure. Figure 4As shown, device 400 may include processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit system capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or “CPU”), graphics processing units (or “GPU”), neural processing units (“NPU”), microcontroller units (“MCU”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), application-specific integrated circuits (ASICs), and any combination thereof. In some embodiments, processor 402 may also be a group of processors grouped into individual logic components. For example, such as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b and processor 402n.

[0054] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, such as Figure 4 As shown, the stored data may include program instructions (e.g., program instructions for implementing stages of the process in the framework 200 or framework 300) and data for processing (e.g., video sequences, video bitstreams, or video streams). The processor 402 can access the program instructions and the data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped into a single logical component. Figure 4 (Not shown in the text).

[0055] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), etc.

[0056] For ease of explanation and to avoid ambiguity, processors 402a-402n and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be wholly or partially integrated into any other component of the device 400.

[0057] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transceivers, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0058] In some embodiments, the device 400 may optionally include a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4 As shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video files), etc.

[0059] It should be noted that the video codec (e.g., the codec that executes the processes in frame 200 or frame 300) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all of the process stages in frame 200 or frame 300 can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all of the process stages in frame 200 or frame 300 can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0060] Next, generative video coding consistent with embodiments of this disclosure will be described. Figure 5 This is a schematic diagram illustrating the basic framework 500 of the depth-based video generative compression scheme based on a first-order motion model (FOMM) according to some embodiments of the present disclosure.

[0061] With the advent of deep generative models, including Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs), facial video compression can achieve significant performance improvements. Although various algorithms can reconstruct frames using a small number of facial parameters through the powerful rendering capabilities of deep generative models, certain head pose movements and facial expression movements still cannot be accurately rendered compared to the original moving video.

[0062] exist Figure 5 In this embodiment, the FOMM deforms the reference source frame to follow the motion of the driving video. This method is applicable to various types of video and can be used for facial animation applications. The FOMM employs an encoder-decoder architecture and integrates motion transfer module components.

[0063] like Figure 5 As shown, in the frame 500, the encoder 510 is configured to encode the source frame 512 via an image / video compression method such as HEVC / VVC or JPEG / BPG. In some embodiments, the VVC is used to compress the source frame 512 to obtain a bitstream 522.

[0064] like Figure 5 As shown, in the frame 500, the encoder 510 is configured to use a keypoint decoder (e.g., a face detector) 516 to extract motion representation 518 of the driving frame 514, which can be encoded using arithmetic coding to obtain a bitstream 524. The decoder 530 is configured to decode the bitstreams 522 and 524 to obtain the reconstructed source frame 532 and the reconstructed motion representation 534. Then, in the motion module 540 within the decoder 530, a keypoint decoder (e.g., a face detector) 542 is configured to process the reconstructed source frame 532 to obtain source frame keypoint information 544. Similarly, the reconstructed motion representation 534 can be used to obtain driving frame keypoint information 546. By combining the source frame keypoint information 544 and the driving frame keypoint information 546, keypoints and local affine transformations 548 can be obtained, including each keypoint and a Jacobian matrix computed at each keypoint location. Then, the dense motion network 549 is configured to obtain and output the dense motion field and occlusion map 550 based on the key points and local affine transformation 548.

[0065] For example, an isovariant loss can be used to learn a keypoint extractor without explicit labels. This keypoint extractor computes two sets of ten learned keypoints for each of the source and driving frames. The learned keypoints are derived from a 64×64 channel-size feature map using a Gaussian mapping function, thus each corresponding keypoint can represent feature information from different channels. In the above description, each keypoint is represented by coordinates (x, y), which represents the most important information in the feature map.

[0066] As another example, the dense motion network 549 can use the marker points and the reconstructed source frame 532 to generate the dense motion field and the occlusion map 550.

[0067] Finally, the generation module 560 is configured to output an image 570 of the object. For example, the generation module 560 may include a neural network configured to warp the resulting feature map using the dense motion field through a differentiable grid sampling operation, and then multiply the warped map with an occlusion map. For example, the generation module 560 may include a generative neural network, also referred to as a "generator" in VAE systems, GAN-based systems, or generative face video compression (GFVC) systems. The generator can be trained to reconstruct video frames using the reconstructed source frame 532 and the dense motion field and occlusion map 550. In some embodiments, the dense motion network 549 can convert one or more decoded facial representation parameters into one or more dense motion streams, each of which has a general format that satisfies the requirements of a general generative model of the generator in the generation module 560. In some embodiments, the dense motion network 549 also converts the one or more decoded facial representation parameters into one or more occlusion maps, each of which has a general format that satisfies the requirements of the general generative model of the generator in the generation module 560. Figure 5 In the FOMM encoder-decoder architecture, the decoder 530 is configured to generate image 570 from the distorted image via the generative neural network trained in the generation module 560.

[0068] Figure 6 This is a schematic diagram illustrating an encoder 600 of a depth-based video generative compression scheme based on Compact Feature Temporal Evolution (CFTE) according to some embodiments of the present disclosure. Figure 7 This is a schematic diagram illustrating a decoder 700 of a depth-based video generative compression scheme based on Compact Feature Temporal Evolution (CFTE) according to some embodiments of the present disclosure. Figure 6-7 The framework in the code follows the encoder-decoder architecture described above.

[0069] like Figure 6 As shown, at the encoder end, the encoder 600 of the compression framework includes a VVC encoder 610, a feature extractor 620, and a feature encoding module 630. The feature extractor 620 can function as a compact key-map detector, and the feature encoding module 630 can function as a context-based entropy encoding module. The VVC encoder 610 is configured to compress the keyframe KF1. The feature extractor 620 is configured to extract compact human features from the keyframe KF1 and other inter-frame frames IF1-IFn. The feature encoding module 630 is configured to compress the inter-frame prediction residuals of the compact human features. First, the keyframe KF1, which represents human texture, is compressed using the VVC encoder 610. Through the compact feature extractor 620, the keyframe KF1 can be represented by a compact feature matrix 640 of size 1×4×4 (e.g., a quantized keymap), and each of the subsequent inter-frame frames IF1-IFn can be represented by a compact feature matrix 650 of size 1×4×4 (e.g., a quantized keymap). In some embodiments, the sizes of the compact feature matrices 640 and 650 are not fixed, and the number of feature parameters can be increased or decreased according to specific bit consumption requirements. These extracted features are then inter-frame predicted and quantized into a residual matrix 660 (e.g., an inter-frame prediction keymap). The feature encoding module 630 is then configured to encode the residual entropy into the bitstream 670.

[0070] In addition, such as Figure 7 As shown, at the decoder 700, the compression framework includes a VVC decoder 710, a feature extractor 720, a feature decoding module 730, a sparse and dense motion module 740, and a generation module 750. The feature extractor 720 may be a compact keymap detector, and the feature decoding module 730 may be a context-based entropy decoding module.

[0071] The VVC decoder 710 is configured to obtain the reconstructed keyframe KF1' based on the received bitstream 670. The feature extractor 720 is configured to extract the compact human features from the reconstructed keyframe KF1' to obtain the reconstructed compact feature matrix 760 (e.g., a quantized keymap) of size 1×4×4. Thus, during the generation of the video, the decoded keyframe KF1' from the bitstream 670 can be further characterized in the form of features through compact feature extraction. The feature decoding module 730 is configured to output the reconstructed residual matrix 770 (e.g., an inter-frame prediction keymap), which includes the reconstructed inter-frame prediction residuals of the compact human features. Thus, when reconstructing the compact features, a reconstructed compact feature matrix 780 (e.g., a compensated keymap) of size 1×4×4 can be obtained for each of the inter-frames IF1-IFn through entropy decoding and compensation. Subsequently, given the features from the keyframes and the inter-frames, the sparse and dense motion module 740 is configured to compute the associated sparse motion field and facilitate the generation of pixel-level dense motion maps and occlusion maps. Finally, based on the deep generative model, the generation module 750 is configured to use the decoded keyframes, pixel-level dense motion maps, and occlusion maps with implicit motion field properties to generate a final video 790 with accurate appearance, pose, and expression. Therefore, the final video 790 can be generated by fully utilizing the reconstructed features and decoded keyframes.

[0072] While the aforementioned generative video compression techniques can achieve satisfactory rate-distortion (RD) performance, they may have associated drawbacks and challenges that limit further performance improvements and practical applications. For example, the flexibility of the generative video codec may be limited by feature extraction distortion at a fixed feature size, thus preventing it from handling inputs of different resolutions.

[0073] In some embodiments of this disclosure, solutions are provided to address one or more of the aforementioned problems and challenges associated with generative video compression.

[0074] In some embodiments, to further enhance the adaptability and flexibility of generative video coding, a scalable resolution neural network can dynamically adjust its network width and depth to adapt to inputs at different resolutions. In deep image coding, the input image is transformed into a latent space for entropy coding. These latent features are low-level features that contain subtle differences in the image and remain compatible across different sizes. However, in generative video coding, features and motion are high-level information, so a model is typically trained and inferred for a single resolution. In some embodiments of this disclosure, the proposed framework uses a more dynamic network structure to achieve more general multi-resolution scalability.

[0075] In some embodiments, a scalable resolution neural network can be used for both foreground and background generation. For lack of generality, the most probable input resolution is denoted as r, and the number of supported resolutions is denoted as N. s According to equation (1), the disclosed scalable resolution neural network is designed to support multiple resolutions with a downsampling factor k:

[0076] Assume there is a downsampling factor s between the motion m and the input image. During generation, the number (depth) of encoder or decoder blocks in the neural network is N. B =log2s, to match the dimensions of the motion. To handle all resolutions, generate a width of N. s And not greater than N B .

[0077] In the encoder section, the keyframe reconstruction is downsampled to features of the same size as the foreground motion, and these features are distorted by the foreground motion. Based on the desired output resolution r... i According to equation (2), in all N B There are blocks One upsampled block: in, Let F represent the reconstructed keyframe, m represent the estimated motion flow, * represent the warp operation, d represent the downsampled block, and g represent the ordinary decoder block that preserves the feature size. Then, after each decoder block, according to equation (3), the feature is compared with the warped feature F from the corresponding block of the encoder portion having the corresponding feature size. -i Weighted summation: F i =b i (F i-1 )·(1-occ)+(F -i *m)·occ (3), Among them, b i The block represented in i <n u If the feature size is not specified, it will be an upsampled block; otherwise, it will be a regular decoder block that retains the specified feature size. Finally, according to equation (4), the reconstructed inter-frame can be obtained by activating the last decoder feature with the activation function σ:

[0078] Figure 8 This is a schematic diagram illustrating an example network structure 800 of a scalable resolution neural network 830 according to some embodiments of the present disclosure. In some embodiments, the neural network 830 is a trained generator, i.e., the neural network 830 may be trained, for example, in conjunction with a discriminator, using a deep generative model. In some embodiments, the neural network 830 may be a deep neural network, particularly a deep generative network with strong inference capabilities to reconstruct realistic images. The trained generator may take dense motion streams (e.g., motion graph 820) and occlusion maps (e.g., occlusion map 810) as input and reconstruct the frames (e.g., reconstructed frame 840), such as facial frames. Figure 8 The network structure 800 described in the document provides an example, in which N s =N B =3. In Figure 8 In this embodiment, the scalable resolution neural network 830 includes multiple blocks 831 that maintain the feature size, multiple downsampling blocks 833, and multiple upsampling blocks 835, where "↑" represents an upsampling block, "↓" represents a downsampling block, "→" represents a block that maintains the feature size, "w" represents a distortion operation 837 using a motion map 820, and "×" represents a masking operation 839 using an occlusion map 810. The number of decoder blocks in the neural network 830 is log2s, where s is the downsampling factor between the motion information and the reconstructed keyframe. In some embodiments, the network width of the neural network 830 is less than or equal to the network depth of the neural network 830 (i.e., the number of encoder or decoder blocks).

[0079] The network structure 800 can be automatically initialized according to the depth and width settings. For example, Figure 8 The network depth is set to 3. Additionally, assuming the network width is configured to 1, only modules with solid line outlines will be initialized. Assuming the network width is configured to 3, then... Figure 8All modules will be initialized. Accordingly, the scalable resolution neural network 830 can dynamically adapt its network depth and width to inputs of different resolutions and is configured to obtain the reconstructed frame 840 and produce a final video with accurate appearance, pose, and expression. As described above, the scalable resolution neural network 830 can be configured to support multiple resolutions R0-R N-1 , where R i From R / k i Defined as R, where R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions. The number of supported resolutions can be achieved by adjusting the number of encoder or decoder blocks and the network width. Accordingly, inputs of different resolutions will pass through different paths in the scalable resolution neural network 830 to maintain their resolution for reconstruction.

[0080] Figure 9 This is a schematic diagram illustrating an example generative video coding system 900 with a scalable resolution neural network 938 according to some embodiments of the present disclosure. Figure 9 As shown, the generative video coding system 900 includes an encoder 910 and a decoder 920. Both the encoder 910 and the decoder 920 can be implemented as devices (e.g., Figure 4 One or more software or hardware components of the device 400 in the middle.

[0081] Return to reference Figure 9 In some embodiments, the encoder 910 includes an encoding module 912 (e.g., using a VVC codec) and a feature factor decomposition module 914. The encoding module 912 is configured to encode and output an image bitstream 922, which includes encoded information for a keyframe KF1 of the video sequence. The feature factor decomposition module 914 is configured to obtain extracted features and encode a feature bitstream 924, which includes encoded information for extracted features of one or more inter-frame frames IF1-IFn of the video sequence.

[0082] In some embodiments, the decoder 920 includes a decoding module 932, a feature factorization module 934, a motion predictor 936, and a scalable resolution neural network 938. In some embodiments, the scalable resolution neural network 938 may be composed of... Figure 8The scalable resolution neural network 830 shown is implemented as described. The decoding module 932 (e.g., using a VVC codec) is configured to decode the image bitstream 922 to reconstruct the keyframes and obtain the reconstructed keyframe KF1'. The feature factorization module 934 is configured to obtain extracted features of the reconstructed keyframe KF1'. The motion predictor 936 is configured to obtain motion information (e.g., pixel-level dense motion map 820) and occlusion information (e.g., occlusion map 810) based on the extracted features of the reconstructed keyframe KF1' and the extracted features of the one or more inter-frame frames IF1-IFn. Accordingly, the scalable resolution neural network 938 can be configured, based on the deep generative model, using the reconstructed keyframe KF1', the pixel-level dense motion map 820, and the occlusion map 810 with implicit motion field characteristics, to obtain the reconstructed frames RF1-RFn and generate the final video with accurate appearance, pose, and expression.

[0083] Figure 10 This is a flowchart of an example method 1000 for decoding a bitstream according to some embodiments of the present disclosure. The method 1000 can be executed by a decoder to decode a video bitstream. For example, the decoder can be implemented for decoding the bitstream (e.g., Figure 9 An apparatus (e.g., containing image bitstream 922 and feature bitstream 924) to reconstruct video frames or video sequences of the bitstream. Figure 4 One or more software or hardware components of the device 400. For example, a processor (e.g., Figure 4 The processor 402) can execute the method 1000. Figure 10 As shown, the method 1000 includes the following steps 1010-1050.

[0084] In step 1010, the decoder receives an image bitstream associated with the video sequence (e.g., Figure 9 Image bitstream 922) and feature bitstream 924 (e.g., Figure 9 The bitstream received from the encoder includes encoded information associated with a keyframe in the video sequence and one or more inter-frames following the keyframe.

[0085] In steps 1020-1080, after receiving the bitstream, the decoder can decode the bitstream to output a video sequence. In step 1020, after receiving the bitstream, the decoder can decode the image bitstream to reconstruct keyframes of the video sequence and obtain extracted features of the reconstructed keyframes.

[0086] In step 1030, after receiving the bitstream, the decoder can decode the feature bitstream to obtain extracted features of one or more inter-frame frames of the video sequence.

[0087] In step 1040, based on the extracted features of the reconstructed keyframes and the extracted features of the one or more inter-frame frames, the decoder can obtain motion information (e.g., Figure 9 Pixel-level dense motion graph 820) and occlusion information (e.g., Figure 9 (Occlusion diagram 810 in the middle).

[0088] In step 1050, based on the motion information and occlusion information, the decoder can utilize a neural network (e.g., Figure 9 The scalable resolution neural network 938 in the decoder resamples the reconstructed keyframes. The network width and depth of the neural network used in step 1050 are adjusted according to the input resolution. Then, in step 1060, the decoder reconstructs the video sequence based on the resampled reconstructed keyframes.

[0089] In some embodiments, in steps 1050 and 1060, the decoder may downsample the reconstructed keyframe to obtain downsampled features with the same size as the foreground motion, distort the downsampled features by the foreground motion to obtain distorted features, and generate a weighted sum of the distorted features for obtaining one or more inter-frame frames for reconstruction.

[0090] In some embodiments, in step 1050, the decoder can access one or more downsampling blocks of the neural network (e.g., Figure 8 Downsampling is performed by downsampling block 833 in the neural network, and upsampling is performed by one or more upsampling blocks in the neural network (e.g., Figure 8 Upsampling is performed by the upsampling block 835 in the neural network. In some embodiments, the neural network can dynamically adjust the network width and the network depth to adapt to inputs of different resolutions, so that inputs of different resolutions can pass through different paths in the neural network to maintain the required resolution for reconstruction.

[0091] Figure 11 This is a flowchart of an example method 1100 for encoding a bitstream according to some embodiments of the present disclosure. The method 1100 can be performed by an encoder to encode a video bitstream. For example, the encoder can be implemented for encoding the bitstream (e.g., Figure 9 The image bitstream 922 and feature bitstream 924 in the image bitstream 922 are used as means for reconstructing video frames or video sequences (e.g., Figure 4One or more software or hardware components of the device 400. For example, a processor (e.g., Figure 4 The processor 402) can execute the method 1100. Figure 11 As shown, the method 1100 includes the following steps 1110, 1120 and 1130.

[0092] In step 1110, the encoder receives a video sequence having a keyframe (e.g., Figure 9 The keyframe KF1 in the frame and one or more inter-frames following the keyframe (e.g., Figure 9 The keyframes and the one or more inter-frames are classified.

[0093] In step 1120, the encoder can process the image bitstream (e.g., Figure 9 The image bitstream (922) is encoded, and the image bitstream contains encoded information of the keyframe. The image bitstream can be decoded to reconstruct the keyframe. In some embodiments, the encoder includes an encoding module that encodes and outputs the image bitstream using a VVC codec.

[0094] In step 1130, the encoder can process the feature bitstream (e.g., Figure 9 The encoder encodes a feature bitstream (924) having encoded information for extracting features for the one or more inter-frame frames. In some embodiments, the encoder may use a feature factor decomposition module to obtain extracted features and encode the feature bitstream, which includes encoded information for extracting features for the one or more inter-frame frames.

[0095] During decoding, dense motion and occlusion information are generated using the features of the reconstructed keyframes and the features of one or more inter-frames encoded in the feature bitstream, for the neural network to resample the reconstructed keyframes. The neural network reconstructs the video sequence by adjusting the network width and depth according to the input resolution. The neural network is configured to support multiple resolutions R0, ... and R... N-1 , where R i From R / k i By definition, R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions.

[0096] In some embodiments, the neural network is configured to dynamically adjust the network width and the network depth to adapt to inputs of different resolutions, and the network width of the neural network is less than or equal to the network depth. For example, the neural network may include log2s encoder or decoder blocks, where s is a downsampling factor between the motion information and the reconstructed keyframes. The reconstructed keyframes can be resampled by performing downsampling by one or more downsampling blocks of the neural network and upsampling by one or more upsampling blocks of the neural network. Accordingly, the reconstructed keyframes can be downsampled to obtain downsampled features with the same size as the foreground motion. The downsampled features are distorted by the foreground motion to obtain distorted features, and a weighted sum of the distorted features can be generated and used to obtain one or more reconstructed inter-frames.

[0097] In some embodiments, a non-transitory computer-readable storage medium is also provided on which an image bitstream and a feature bitstream are stored. The image bitstream and the feature bitstream can be encoded and decoded according to a disclosed scalable resolution neural network for generative video compression.

[0098] As described above, an image bitstream and a feature bitstream can be generated based on a video sequence. The image bitstream includes encoded information for reconstructing keyframes of the video sequence and obtaining extracted features of the reconstructed keyframes, and the feature bitstream includes encoded information for obtaining extracted features of one or more inter-frame frames of the video sequence. Accordingly, based on the input video sequence, the image bitstream and the feature bitstream can be generated and stored in at least one non-transitory computer-readable storage medium. The video sequence will be reconstructed by a neural network that resamples the reconstructed keyframes, and the neural network adjusts its width and depth according to the input resolution.

[0099] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (e.g., the disclosed encoder and decoder) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAMs, caches, registers, any other memory chips or cassette tapes, and their networking versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0100] It should be noted that the relational terms such as "first," "second," etc., used in this document are only used to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be identical in meaning and open-ended, because one or more items following any of these words does not imply an exhaustive list of such one or more items, nor does it imply limitation to only one or more listed items.

[0101] As used herein, unless otherwise specified, the term "or" covers all possible combinations unless impractical. For example, if it is specified that a database may include A or B, then unless otherwise specified or impractical, the database may include A, or B, or A and B. As a second example, if it is specified that a database may include A, B, or C, then unless otherwise specified or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0102] The embodiments may be further described using the following terms: 1. A video decoding method, comprising: Decode the image bitstream associated with the video sequence to reconstruct the keyframes of the video sequence and obtain the extracted features of the reconstructed keyframes; Decode the feature bitstream associated with the video sequence to obtain extracted features of one or more inter-frame frames of the video sequence; Based on the extracted features of the reconstructed keyframes and the extracted features of one or more inter-frame frames, motion information and occlusion information are obtained. Based on the motion and occlusion information, the reconstructed keyframes are resampled by a neural network; and Based on the resampled keyframes, the video sequence is reconstructed by the neural network, wherein the network width and network depth of the neural network are adjusted according to the input resolution. 2. The video decoding method according to Clause 1, wherein resampling the reconstructed keyframes includes: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; and Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames. 3. The video decoding method according to clause 1 or 2 further includes: The network width and network depth of the neural network are dynamically adjusted to adapt to inputs of different resolutions. 4. The video decoding method according to any one of clauses 1 to 3, wherein the neural network is configured to support multiple resolutions R0, ... and R... N-1 Where Ri is composed of R / k i By definition, R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions. 5. The video decoding method according to any one of clauses 1 to 4, wherein the neural network comprises log2s decoder blocks, and s is a downsampling factor between the motion information and the reconstructed keyframe. 6. The video decoding method according to any one of clauses 1 to 5, wherein the network width of the neural network is less than or equal to the network depth of the neural network. 7. The video decoding method according to any one of clauses 1 to 6, wherein resampling the reconstructed keyframes comprises: Downsampling is performed by one or more downsampling blocks of the neural network; and Upsampling is performed by one or more upsampling blocks of the neural network. 8. A video encoding method, comprising: Encoding an image bitstream, the image bitstream including: encoded information for keyframes of a video sequence, wherein the image bitstream can be decoded to reconstruct the keyframes; and The feature bitstream is encoded, the feature bitstream including encoded information for extracting features from one or more inter-frame frames of the video sequence. Specifically, dense motion information and occlusion information are generated by using the features of the reconstructed keyframes and the features of one or more inter-frames encoded in the feature bitstream, so that the neural network can resample the reconstructed keyframes. The neural network reconstructs the video sequence by adjusting the network width and depth according to the input resolution. 9. The video coding method according to Clause 8, wherein the reconstructed keyframes are resampled by the following steps: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; and Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames. 10. The video coding method according to clause 8 or 9, wherein the neural network is configured to dynamically adjust the network width and the network depth to adapt to inputs of different resolutions of the neural network. 11. The video coding method according to any one of clauses 8 to 10, wherein the neural network is configured to support multiple resolutions R0, ... and R... N-1 Where Ri is composed of R / k i By definition, R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions. 12. The video coding method according to any one of clauses 8 to 11, wherein the neural network comprises log2s blocks, where s is a downsampling factor between the motion information and the reconstructed keyframes. 13. The video coding method according to any one of clauses 8 to 12, wherein the network width of the neural network is less than or equal to the network depth of the neural network. 14. The video coding method according to any one of clauses 8 to 13, wherein the reconstructed keyframe is resampled by performing downsampling by one or more downsampling blocks of the neural network and upsampling by one or more upsampling blocks of the neural network. 15. A method for storing an image bitstream and a feature bitstream, the method comprising: An image bitstream and a feature bitstream are generated based on a video sequence. The image bitstream includes encoded information for reconstructing keyframes of the video sequence and obtaining extracted features from the reconstructed keyframes. The feature bitstream includes encoded information for obtaining extracted features from one or more inter-frame frames of the video sequence. The image bitstream and the feature bitstream are stored in at least one non-transitory computer-readable medium. The video sequence is reconstructed by a neural network that resamples the reconstructed keyframes, and the neural network adjusts its width and depth according to the input resolution. 16. The method according to Clause 15, wherein the reconstructed keyframe is resampled by the following steps: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; and Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames. 17. The method according to Clause 15 or 16, wherein the neural network is configured to dynamically adjust the network width and the network depth to adapt to inputs of different resolutions of the neural network. 18. The method according to any one of clauses 15 to 17, wherein the neural network is configured to support multiple resolutions R0, ..., R Ns-1 , where R i From R / k i Defined as follows: R is the maximum input resolution, k is the downsampling factor, and N... s It is the number of resolutions. 19. The method according to any one of clauses 15 to 18, wherein the reconstructed keyframe is resampled by the following steps: Downsampling is performed by one or more downsampling blocks of the neural network; and Upsampling is performed by one or more upsampling blocks of the neural network. 20. The method according to any one of clauses 15 to 19, wherein the network width of the neural network is less than or equal to the network depth of the neural network.

[0103] It should be understood that the above embodiments can be implemented in hardware, software (program code), or a combination of hardware and software. If implemented in software, it can be stored in the aforementioned computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented in hardware, software, or a combination of hardware and software. Those skilled in the art should also understand that multiple modules / units described above can be combined into one module / unit, and each module / unit described above can be further divided into multiple sub-modules / sub-units.

[0104] In the foregoing specification, numerous specific details have been described with reference to embodiments, which may vary depending on the implementation. Certain adjustments and modifications may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the appended claims. The sequence of steps shown in the figures is also to be considered for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, those skilled in the art will understand that these steps may be performed in a different order when implementing the same method.

[0105] Exemplary embodiments have been disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments, and the embodiments described in this disclosure can be freely combined. Therefore, although specific terminology has been used, it is used in a general and descriptive sense only and is not intended to be limiting.

Claims

1. A video decoding method, comprising: Decode the image bitstream associated with the video sequence to reconstruct the keyframes of the video sequence and obtain the extracted features of the reconstructed keyframes; Decode the feature bitstream associated with the video sequence to obtain extracted features of one or more inter-frame frames of the video sequence; Based on the extracted features of the reconstructed keyframes and the extracted features of one or more inter-frame frames, motion information and occlusion information are obtained. Based on the motion and occlusion information, the reconstructed keyframes are resampled by a neural network; as well as Based on the resampled keyframes, the video sequence is reconstructed by the neural network, wherein the network width and network depth of the neural network are adjusted according to the input resolution.

2. The video decoding method according to claim 1, wherein, Resampling the reconstructed keyframes includes: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; and Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames.

3. The video decoding method according to claim 1 further includes: The network width and network depth of the neural network are dynamically adjusted to adapt to inputs of different resolutions.

4. The video decoding method according to claim 1, wherein, The neural network is configured to support multiple resolutions R0, ... and R1. N-1 , where R i From R / k i By definition, R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions.

5. The video decoding method according to claim 1, wherein, The neural network comprises log2s decoder blocks, where s is the downsampling factor between the motion information and the reconstructed keyframes.

6. The video decoding method according to claim 1, wherein, The network width of the neural network is less than or equal to the network depth of the neural network.

7. The video decoding method according to claim 1, wherein, The resampling of the reconstructed keyframes includes: Downsampling is performed by one or more downsampling blocks of the neural network; and Upsampling is performed by one or more upsampling blocks of the neural network.

8. A video encoding method, comprising: Encoding an image bitstream, the image bitstream including: encoded information for keyframes of a video sequence, wherein the image bitstream can be decoded to reconstruct the keyframes; and The feature bitstream is encoded, the feature bitstream including: encoded information for extracting features from one or more inter-frame frames of the video sequence. Specifically, dense motion information and occlusion information are generated by using the features of the reconstructed keyframes and the features of one or more inter-frames encoded in the feature bitstream, so that the neural network can resample the reconstructed keyframes. The neural network reconstructs the video sequence by adjusting the network width and depth according to the input resolution.

9. The video encoding method according to claim 8, wherein, The reconstructed keyframes are resampled using the following steps: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; as well as Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames.

10. The video encoding method according to claim 8, wherein, The neural network is configured to dynamically adjust the network width and the network depth to adapt to inputs of different resolutions.

11. The video encoding method according to claim 8, wherein, The neural network is configured to support multiple resolutions R0, ... and R1. N-1 , where R i From R / k i By definition, R is the maximum input resolution, k is the downsampling factor, and N is the number of resolutions.

12. The video encoding method according to claim 8, wherein, The neural network comprises log2s blocks, where s is the downsampling factor between the motion information and the reconstructed keyframes.

13. The video encoding method according to claim 8, wherein, The network width of the neural network is less than or equal to the network depth of the neural network.

14. The video encoding method according to claim 8, wherein, The reconstructed keyframes are resampled by performing downsampling by one or more downsampling blocks of the neural network and upsampling by one or more upsampling blocks of the neural network.

15. A method for storing an image bitstream and a feature bitstream, the method comprising: An image bitstream and a feature bitstream are generated based on a video sequence. The image bitstream includes encoded information for reconstructing keyframes of the video sequence and obtaining extracted features from the reconstructed keyframes. The feature bitstream includes encoded information for obtaining extracted features from one or more inter-frame frames of the video sequence. The image bitstream and the feature bitstream are stored in at least one non-transitory computer-readable medium. The video sequence is reconstructed by a neural network that resamples the reconstructed keyframes, and the neural network adjusts its width and depth according to the input resolution.

16. The method according to claim 15, wherein, The reconstructed keyframes are resampled using the following steps: The reconstructed keyframes are downsampled to obtain downsampled features with the same motion size as the foreground. The downsampled features are distorted by the foreground motion to obtain distorted features; as well as Generate a weighted sum of the distortion features, wherein the weighted sum is used to obtain one or more reconstructed inter-frame frames.

17. The method according to claim 15, wherein, The neural network is configured to dynamically adjust the network width and the network depth to adapt to inputs of different resolutions.

18. The method according to claim 15, wherein, The neural network is configured to support multiple resolutions R0, ..., R1 Ns-1 , where R i From R / k i Defined as follows: R is the maximum input resolution, k is the downsampling factor, and N... s It is the number of resolutions.

19. The method according to claim 15, wherein, The reconstructed keyframes are resampled using the following steps: Downsampling is performed by one or more downsampling blocks of the neural network; and Upsampling is performed by one or more upsampling blocks of the neural network.

20. The method of claim 15, wherein, The network width of the neural network is less than or equal to the network depth of the neural network.

21. A non-transitory computer-readable storage medium storing an instruction set, an image bitstream, and a feature bitstream thereon, the instruction set being executable by one or more processors in a method to decode the image bitstream and the feature bitstream, the method comprising: Decode the image bitstream associated with the video sequence to reconstruct the keyframes of the video sequence and obtain the extracted features of the reconstructed keyframes; Decode the feature bitstream associated with the video sequence to obtain extracted features of one or more inter-frame frames of the video sequence; Based on the extracted features of the reconstructed keyframes and the extracted features of one or more inter-frame frames, motion information and occlusion information are obtained. Based on the motion and occlusion information, the reconstructed keyframes are resampled by a neural network; as well as Based on the resampled keyframes, the video sequence is reconstructed by the neural network, wherein the network width and network depth of the neural network are adjusted according to the input resolution.