Generative video coding method based on implicit inter-frame alignment
Through the generative video encoding method of implicit inter-frame alignment, the adaptive spatiotemporal importance encoder and U-Net network are used to solve the inter-frame feature modeling and key information retention problems of video encoding at low code rates, and improve the perceived quality and compression efficiency of videos.
Patent Information
- Application Number
- CN202510504240.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
Existing video encoding methods are difficult to balance perception quality and compression efficiency under low code rate conditions, and there are problems such as insufficient inter-frame feature modeling and insufficient retention of key information, resulting in a decline in visual perception quality.
The generative video encoding method based on implicit inter-frame alignment is adopted, and the spatial and temporal characteristics of video frames are intelligently modeled through an adaptive spatiotemporal importance encoder, combined with implicit inter-frame alignment strategy and perceived loss function, dynamically allocate the code rate, and use the U-Net network for information fusion and reconstruction.
It improves the perceived quality of video frames, reduces the computational complexity and bit overhead, enhances the inter-frame feature modeling capabilities, and optimizes the video reconstruction effect in low-code rate environments.
Smart Images

Figure CN120434385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video coding, and in particular to a generative video coding method based on implicit inter-frame alignment. Background Art
[0002] Currently, massive amounts of video data in complex environments across a wide range of scenarios are placing enormous pressure on limited transmission bandwidth and storage devices. Traditional video coding standards, such as H.264 / AVC and H.265 / HEVC, use motion estimation and motion compensation to remove spatial redundancy within video frames and temporal redundancy between frames, alleviating the pressure on video storage and transmission to a certain extent. However, the lower the bitrate, the more pronounced the blocking artifacts in the reconstructed video, seriously affecting the visual quality of the video. Furthermore, traditional video coding frameworks rely primarily on complex, manually designed transformations and coding strategies, making them difficult to adapt to the dynamic characteristics of different scenes and content, especially in environments with high compression ratios and low bitrates.
[0003] Deep learning-based video coding methods have made significant progress in recent years. However, problems such as a mismatch between optimization objectives and perceived quality, and biased training data distribution have reduced visual quality at very low bitrates. Generative coding effectively improves texture and structure restoration capabilities at low bitrates by learning from data distribution, alleviating the blurring artifacts associated with traditional deep video compression to some extent. However, existing research still faces two major bottlenecks: insufficient temporal correlation modeling and the lack of inter-frame feature association; and the lack of a dynamic bit allocation mechanism, which hinders the adaptive extraction of key information.
[0004] Furthermore, existing video coding methods often struggle to balance perceived quality and compression efficiency at low bitrates. On the one hand, existing methods suffer from a mismatch between their optimization objectives and perceived quality, leading to a decline in visual quality at extremely low bitrates. On the other hand, biased training data distribution also impacts the generalization capabilities of these methods. Therefore, improving perceived quality at low bitrates while simultaneously enhancing inter-frame feature modeling and preserving key information has become a key research topic in the field of intelligent visual information processing.
[0005] In summary, there is at least one of the following technical problems:
[0006] Deep learning-based video coding methods have made significant progress, but they suffer from issues such as a mismatch between optimization objectives and perceived quality, and biased training data distribution, which degrades visual quality at very low bitrates. Generative coding methods can restore video details to a certain extent by learning data distribution, but they still struggle with modeling inter-frame features and retaining key information.
[0007] Existing video coding methods often struggle to balance perceived quality and compression efficiency at low bitrates. On the one hand, the optimization objectives of these methods are mismatched with perceived quality, leading to a decrease in visual quality at very low bitrates. On the other hand, biased training data distribution also affects the generalization capabilities of these methods.
[0008] In summary, how to improve the perceptual quality of video under low bit rate conditions while strengthening the inter-frame feature modeling capability and retaining key information is a technical problem that the present invention needs to solve. Summary of the Invention
[0009] The main purpose of the present invention is to provide a generative video coding method based on implicit inter-frame alignment, so as to solve the problem of the existing deep learning video coding technology that the visual perception quality is degraded at extremely low bit rates due to the mismatch between the optimization goal and the perceptual quality and the deviation in the distribution of training data. The generative coding method can restore the video details to a certain extent by learning the data distribution, but it still has shortcomings in inter-frame feature modeling and key information retention. Existing video coding methods are often difficult to balance the perceptual quality and compression efficiency of the video under low bit rate conditions. On the one hand, the method has the problem of mismatch between the optimization goal and the perceptual quality, resulting in a decrease in visual perception quality at extremely low bit rates; on the other hand, the deviation in the distribution of training data also affects the generalization ability of the method. The present invention needs to solve the technical problem of how to improve the perceptual quality of the video under low bit rate conditions, while strengthening the inter-frame feature modeling capability and retaining key information.
[0010] To achieve the above object, according to one aspect of the present invention, a generative video coding method based on implicit inter-frame alignment is provided, comprising:
[0011] Step 1: The adaptive spatiotemporal importance encoder intelligently models the spatiotemporal characteristics of video frames and dynamically allocates bitrates.
[0012] Step 2: The features of the previous frame are extracted through the implicit inter-frame alignment strategy and then fused with the information of the current frame. The information of the previous frame is transformed and aligned with the current frame through the structure of the U-Net network.
[0013] Step 3: Combine perceptual image patch similarity constraints through perceptual loss function.
[0014] Preferably, the step 1 includes: setting an original frame to be encoded, extracting the original frame through a feature extractor to obtain target features, and using reference frame features as prior information to encode the target features to obtain decoding features.
[0015] Preferably, the decoding feature is expressed as:
[0016]
[0017] where E(·) and D(·) represent the encoder and decoder with four-fold up- and down-sampling convolutional layers and two spatiotemporal modulation units, respectively. represents the reference feature; y i Represents target features; Decoding features.
[0018] Preferably, step 1 also includes: extracting spatiotemporal importance weights from reference features using a spatiotemporal attention network through a spatiotemporal modulation unit, which are used to scale current target features and perform bitrate allocation; wherein the reference features and target features are fused through a Concat operation, and then the spatiotemporal importance weights are obtained through a 3×3 convolutional layer with a ReLU activation function and a Sigmoid function.
[0019] Preferably, the step 1 further comprises: balancing the fusion feature f by using the spatiotemporal importance weights through the spatiotemporal modulation unit i and Conv(f i ) to extract the spatiotemporal features in the target features; the spatiotemporal importance weights and bit rate weights are further combined through the spatiotemporal modulation unit to perform dynamic bit rate allocation to obtain the scaled features.
[0020] Preferably, step 2 includes: performing DDIM inverse diffusion noise addition on the reference feature to make it close to the distribution of Gaussian noise, and replacing the original random noise ε~N(0,1) with the initialization feature as the initial input of the diffusion model.
[0021] Preferably, the step 2 further comprises: aligning the target features with the initialization features between frames and performing denoising optimization using U-Net to obtain high-quality latent space features; the state transition process of the diffusion model It can be regarded as a process of gradually denoising from the initial features to the maximum likelihood features.
[0022] Preferably, the maximum likelihood feature is expressed as:
[0023]
[0024] in, It represents the denoising prediction model based on U-Net, which is the total process of diffusion denoising; t represents the number of diffusion steps; represents the maximum likelihood feature; Represents the initialization feature.
[0025] Preferably, step 3 includes: a loss function setting strategy, which is divided into two stages: the first stage uses the initial frame and a single forward frame to train the model, controlling the reconstruction loss from the initial frame to the forward frame; the second stage considers increasing the number of forward frames to learn the correlation between forward frames and further control the overall loss. The loss function is constructed by combining the bit rate and reconstruction distortion results obtained from the two stages of training to achieve better compression effect.
[0026] Preferably, during the state transfer process, the parameters of the U-Net network remain fixed, and the rate-distortion loss function Expressed as:
[0027]
[0028] Where λ represents the bit rate and the balance coefficient between reconstruction distortion, β represents the mean square error MSE loss and learning-aware patch similarity LPIPS loss The weight coefficient between represents the rate-distortion loss function; y i Represents target features; Represents mean square error MSE loss; represents the learning-perceptual image patch similarity LPIPS loss; x i (i∈{0,1,…n}) represents the input video sequence; Represents the reconstructed video sequence. Under the above loss function setting strategy, The following formula will be obtained:
[0029]
[0030] in, To use the I frame and the previous frame for training, the bit rate, MSE loss, and LPIPS loss of the previous frame are obtained; To use the previous frame and the current frame for training, we can get the bit rate, MSE loss, and LPIPS loss of the current frame; i represents the target feature; x i (i∈{0,1,…n}) represents the input video sequence; Represents the reconstructed video sequence.
[0031] The application of the technical solution of the present invention has the following technical effects:
[0032] An adaptive spatiotemporal importance encoder optimizes the quality of key regions while reducing the bit overhead of less important regions. An implicit inter-frame alignment strategy captures the underlying spatiotemporal relationships between frames, reducing the computational complexity of estimating explicit motion information and computational overhead. A perceptual loss function effectively improves the perceptual quality of video frames.
[0033] By combining a diffusion model with an implicit inter-frame alignment strategy and using decoded features as preconditions for pre-training the latent diffusion model, this method captures latent features between frames, thereby recovering target features and reconstructing video frames. This reduces the redundant information found in traditional motion estimation methods while better recovering detail and structural information between video frames. Furthermore, through an adaptive spatiotemporal importance encoder and a proposed spatiotemporal modulation unit, target features are scaled. A neural network is used to extract the importance weight of each frame, and the bitrate allocation is dynamically adjusted based on factors such as motion complexity and texture complexity, thereby improving reconstruction quality. Furthermore, by combining learning to perceive image block similarity, the method optimizes video visual quality in low-bitrate environments, overcoming the limitations of traditional evaluation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0035] Figure 1 A method flow chart of a generative video coding method based on implicit inter-frame alignment according to the present invention is shown;
[0036] Figure 2 Shown Figure 1 Implementation principle diagram of the generative video coding method based on implicit inter-frame alignment;
[0037] Figure 3 Shown Figure 1 A structural view of the adaptive spatiotemporal importance encoder for implicit inter-frame alignment-based generative video coding in
[15] .
[0038] Figure 4 Shown Figure 1 A view of the spatiotemporal modulation unit structure of the implicit inter-frame alignment-based generative video coding method in
[15] ;
[0039] Figure 5 Shown Figure 1 The implicit inter-frame alignment module structure view of the generative video coding method based on implicit inter-frame alignment;
[0040] Figure 6 Shown Figure 1Objective comparison results of performance indicators of generative video coding methods based on implicit inter-frame alignment in
[15] .
[0041] Figure 7 Shown Figure 1 First subjective comparison results of the implicit inter-frame alignment-based generative video coding method in
[15] ;
[0042] Figure 8 Shown Figure 1 Second subjective comparison results of the implicit inter-frame alignment-based generative video coding method in
[15] ;
[0043] Figure 9 Shown Figure 1 Heat map of the bitrate distribution of detail features after quantization in the implicit inter-frame alignment-based generative video coding method. DETAILED DESCRIPTION
[0044] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0045] like Figures 1 to 9 As shown, an embodiment of the present invention provides a generative video coding method based on implicit inter-frame alignment, including: step 1: intelligently modeling the spatiotemporal characteristics between video frames through an adaptive spatiotemporal importance encoder, and dynamically allocating the bit rate; step 2: extracting the features of the previous frame through an implicit inter-frame alignment strategy and fusing them with the information of the current frame, and transforming the information of the previous frame to align it with the current frame through the structure of the U-Net network; step 3: combining the perceptual image block similarity constraint through a perceptual loss function.
[0046] This patent solves the limitations of explicit optical flow estimation in traditional video coding methods in low bit rate scenarios, such as the blur and artifact problems of reconstructed images, while improving the perceptual quality of the video, enhancing the ability to model inter-frame features and retaining key information.
[0047] This embodiment mainly includes an adaptive spatio-temporal importance-aware codec (ASTC) and an implicit frame alignment module. The adaptive spatio-temporal importance codec can effectively capture the potential spatio-temporal correlations between video frames. By intelligently modeling the spatio-temporal characteristics of video frames, it dynamically adjusts the information encoding strategy to optimize the video compression and reconstruction process. The implicit frame alignment module extracts the features of the previous frame and fuses them with the information of the current frame. Through the U-Net network structure, the previous frame information is transformed and aligned with the current frame to capture the potential spatio-temporal relationship between frames.
[0048] In this embodiment, the processing flow is as follows Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 As shown, the processing steps include the following:
[0049] Step S10: The adaptive spatiotemporal importance encoder uses the information of the original frame to obtain the decoding features. Set the original frame sequence to be encoded as {x t}, obtain the target feature y through feature extraction i , the features of the reference frame As prior information, and the target feature y i After fusion, the decoding features are obtained
[0050] Step S20: Implicit inter-frame alignment module uses the information of the previous frame to learn the spatiotemporal correlation between frames through U-Net to achieve implicit inter-frame alignment. This module mainly consists of three parts: encoder, U-Net and decoder. The input data is mapped to the latent space through the encoder and the forward noise operation is performed to obtain Z T , then extract the condition τ θ , fused with the latent space representation after noise addition, input into the U-Net network for reverse denoising, and then complete the decoding.
[0051] Specifically, the above step S10 includes:
[0052] Total solution The expression is as follows:
[0053]
[0054] where E(·) and D(·) represent the encoder and decoder with four-fold up- and down-sampling convolutional layers and two spatiotemporal modulation units, respectively. represents the reference feature; y i Represents target features; Decoding features.
[0055] In traditional encoders, the coding strategies for different regions are fixed. The spatiotemporal modulation unit (STMU) of the present invention uses a neural network to extract the importance weight of each frame and dynamically adjusts the bit rate distribution according to factors such as motion complexity and texture complexity. The structure is as follows: Figure 4 shown.
[0056] Specifically, STMU uses a spatiotemporal attention network to extract features from the reference Extract spatiotemporal importance weights ω from i , used to scale the current target feature y i And perform code rate allocation, which can be expressed as:
[0057]
[0058] Among them, f i express and y i Concat(·) represents the fusion operation, Conv(·) represents a 3×3 convolution layer with ReLU activation function, and σ(·) represents the Sigmoid function, which is used to normalize the attention weight so that its value is limited to [0,1]. Then, STMU uses the weight ω i Extract target feature y i The spatiotemporal characteristics of It can be expressed as:
[0059]
[0060] Among them, ω i represents the spatiotemporal importance weight, f i express and y i The fusion feature of Conv(·) represents a 3×3 convolution layer with ReLU activation function. After STMU fuses the spatiotemporal attention information, it is i The target features are given different importances, the features of key areas are enhanced, and the features of non-key areas are suppressed. In addition, the module can adaptively adjust the feature expression mode, and through the residual connection, ensure that the feature adjustment does not lead to information loss, thereby improving the stability of the model. STMU further combines the importance weight ω i Dynamic bit rate allocation is performed with the bit rate weight λ to obtain the scaled feature The size is (192, 28, 52), which can be expressed as:
[0061]
[0062] Among them, ω i represents the spatiotemporal importance weight, and λ represents the rate control weight, which determines the bit allocation of the target feature in different regions; represents spatiotemporal features, with a size of (192, 56, 104); Concat(·) represents the fusion operation, Concat(ω i ,λ) combines the spatiotemporal attention weight and the bit rate control factor to adaptively allocate the bit rate. Taking batch=1 as an example, ω i The sizes of and are (4, 1) and (4, 1), respectively. MLP(·) represents a multi-layer perceptron, which is used to learn the rate control strategy. This process ensures that the encoder can dynamically adjust the bit rate based on the importance of spatiotemporal features, improving the compression quality of critical areas while reducing the bit overhead of less important areas.
[0063] Figure 9 The heat map of the bitrate distribution of detail features quantized by the adaptive spatiotemporal importance encoder is presented. In video frames, relatively static areas such as roads and vehicles waiting at traffic lights are relatively common in video sequences and change less. The adaptive spatiotemporal importance encoder allocates fewer bits to represent these areas, making full use of the information provided by the previous frame. For unique and dynamically changing objects such as vehicle trajectories, the adaptive spatiotemporal importance encoder adaptively allocates more bits for encoding. This shows that the adaptive spatiotemporal importance encoder can dynamically adjust the information encoding strategy based on the potential spatiotemporal correlation between video frames. While ensuring that the information quality of key areas is not lost, the overall bitrate is effectively reduced, thereby achieving more efficient video compression and reconstruction.
[0064] Specifically, the above step S20 includes:
[0065] During the decoding process, the decoding features As the initial feature To the target feature y i The maximum likelihood estimation condition of , inter-frame alignment is performed through the implicit diffusion model, and Solve for y i Latent variables of equal dimensions It can be expressed as:
[0066]
[0067] in, represents the initial state of maximizing the likelihood estimate, Represents the reference feature The result of DDIM inverse transformation, T represents the number of steps in the denoising process, represents the number of steps of the DDIM inverse transformation, t=T→1(T=30) represents the number of steps of solving the maximum likelihood estimation, and Represent the mean and variance of the results of the DDIM inverse transformation, y i represents the target feature, represents the decoding feature, Indicates the characteristics of diffusion step t in the denoising process. Each step of the state transition process The calculation is performed through the LDM pre-trained U-Net network, which can be expressed as:
[0068]
[0069] in, and They represent the mean and variance of the fitting during the state transition process respectively. represents the decoding feature, It represents the feature of the t-step diffusion in the denoising process, and T represents the number of steps in the denoising process. With the initial features The state transition process can more accurately predict the mean μ of each step, as it has similarities in structure and semantics but different motion information in the time domain. θ and variance
[0070] In order to reduce the amount of encoded information and not encode motion information while ensuring the reconstruction quality of the video frame, the Implicit Inter-frame Feature Alignment Module (IIFA) is used to decode the features. As a condition for the pre-trained diffusion model - LDM image super-resolution model, the target feature y is recovered i , reconstruct the video frame, the structure is as follows Figure 4 The present invention adopts an implicit inter-frame alignment strategy, aiming to utilize the latent space information and motion information of the reference frame to ensure the reconstruction quality, eliminate the influence of the LDM pre-training model not taking the inter-frame reference information into account during training, and the initial distribution of the maximization likelihood function process is random noise ε~N(0,1).
[0071] In order to effectively align inter-frame information during the diffusion process, the present invention first performs a Perform DDIM inverse diffusion noise to make it close to the distribution of Gaussian noise, and use To replace the original random noise ε~N(0,1) as the initial input of the diffusion model. Solution Used by arrive The recursive formula can be expressed as:
[0072]
[0073] in, represents the initial reference feature, the feature extracted directly from the previous frame, α t-1 and α t Represents the noise weight coefficient of the diffusion model, ω i Indicates the spatiotemporal importance weights used to dynamically adjust the amount of noise, ε θ Represents the output of the diffuse noise prediction model. This process can generate an initialization feature that is more consistent with the distribution of video frames. Reduce the distribution deviation caused by directly using Gaussian noise as the initial input. At the encoding end, the present invention uses the decoding feature and initialization features Perform inter-frame alignment and use U-Net for denoising optimization to obtain high-quality latent space features State transition process of the diffusion model It can be regarded as the initial feature Stepwise denoising to maximum likelihood features The process can be expressed as:
[0074]
[0075] in, Represents the denoising prediction model based on U-Net, which is the total process of diffusion denoising, and t represents the number of diffusion steps. In this process, the initial features Gradually denoise and finally obtain high-quality latent space features Can be used to generate the final decoded video frame; Denotes the decoding feature.
[0076] From the above inter-frame alignment process, it can be seen that the reference feature Provides target features during denoising Therefore, it is crucial to reasonably set the DDIM inverse diffusion step number T' for inter-frame motion modeling. Assuming that the initial features in the diffusion process are Can be represented as a reference feature and noise The sum is:
[0077]
[0078] in, Represents the denoising prediction model based on U-Net, t=T→1(T=30) represents the number of steps to solve the maximum likelihood estimation, It represents the characteristics of the diffusion step t in the denoising process, T represents the number of steps in the denoising process, Represents the reference feature, from the reference feature To the maximum likelihood feature Movement information between It can be expressed as:
[0079]
[0080] in, For sports information, is the maximum likelihood feature, is the reference feature, ε θ is the noise estimation. From this formula, we can see that in the diffusion process, the change of inter-frame motion information is determined by the noise estimation model ε θ Modeling is performed, thereby avoiding the explicit motion estimation process and realizing implicit modeling of inter-frame information.
[0081] Considering the influence of the DDIM inverse transform step number T' on the motion information of the modeling reference feature, the present invention selects an appropriate step number to optimize the inter-frame prediction quality. i With reference features The gap in motion information between the two increases, and the model requires additional motion compensation information, resulting in greater computational overhead. If T' is too small, the target feature y i With reference features The motion information is close, the model can rely more on To reduce coding conditions In order to balance the inter-frame information modeling effect and the bit rate allocation, the present invention sets the number of DDIM inverse transform steps to half the number of diffusion process steps.
[0082] In order to verify the effectiveness of the present invention, the present invention completed a series of experiments.
[0083] The experiments in this invention used the commonly used Vimeo-90K Septuplet dataset to train the encoding model. This dataset contains 89,800 video clips depicting diverse real-world motion across a wide range of scenarios. Each video sequence consists of seven consecutive frames. After preprocessing, 64,612 seven-frame sequences with a resolution of 256×256 were used.
[0084] This experiment selected three classic datasets shot by drones: UAVDT (forest, night, road), ERA (Cycling, Fire, Harvesting), and AU-AIR, with resolutions of 1280×704, 640×640, and 832×448, respectively, and the number of video sequence frames is 32. At the same time, the present invention also selected traditional HEVC datasets, including Class-C (BasketballDrill, BQMall, PartyScene, RaceHorses), Class-D (BasketballPass, BlowingBubbles, BQSquare, RaceHorses), and Class-E (FourPeople, Johnny, KristenAndSara), with resolutions of 832×448, 384×192, and 1280×704, respectively, and the number of video sequence frames is 32.
[0085] This experiment was completed on an Intel(R) Xeon(R) Gold 6426Y CPU platform. Four NVIDIA GeForce RTX 4090 GPUs were used for parallel computing during the training phase (with a batch size of 1 per GPU), and one GPU of the same model was used during the testing phase (with a batch size of 1). The present invention trains the Vimeo-90K Septuplet dataset for four cases of balancing coefficient λ = 1, 8, 256, and 512, wherein when λ = 1 and 512, the learning rate is set to lr = 1×10 –6 , when λ=8, 256, the learning rate is set to lr=1×10 –5 In the testing phase, the above λ parameter configuration is used. All available pre-trained models are evaluated using the LPIPS indicator, and only the model with the best performance is displayed.
[0086] To evaluate the present invention, we compared it with deep learning-based video coding methods, such as the conditionally guided video coding methods DCVC, DCVC-TCM, DCVC-HEM, DCVC-DC, and DCVC-FM. Four different reconstruction qualities were selected for comparison for each comparison method. Furthermore, RD curves were used to demonstrate the objective performance of the present invention and the comparison methods.
[0087] according to Figure 6 The objective results shown are that the horizontal axis of the image has the number of bits required for encoding per pixel (bit perpixel, bpp), and the vertical axis is the learned perceptual image block similarity LPIPS. Compared with other video coding methods, the LPIPS value of the present invention is significantly lower than that of the comparison method in each data set, and as the bpp increases, the LPIPS value of the present invention decreases faster. Not only does it perform well on the traditional video coding data set HEVC, but it also obtains high-quality video frames on some of the now widely covered aerial videos (UAVDT, ERA, AU-AIR). It can be seen that the perceptual quality of the reconstructed image of the present invention is closer to the original image, and has good adaptability under different bit rates. While ensuring the video quality, it effectively reduces the encoding bit rate, meeting the requirements for video compression quality and bit rate when a large amount of shooting and transmission are required in various scenarios.
[0088] according to Figure 7The subjective results shown show that the present invention achieves better subjective reconstruction results than the comparison methods. For the first group of examples, there are a large number of blurring artifacts in the reconstructed frames of the comparison methods DCVC and others, and the texture on the floor cannot be clearly reconstructed. However, the present invention successfully reconstructs the nails on the floor at a similar compression ratio, and the texture of the floor is also closer to the original video frame. For the second group of examples, the smoke at the fire scene reconstructed by the comparison methods DCVC and others is blurred, while the present invention successfully reconstructs the details of the branches obscured by the smoke at a larger compression ratio, and more accurately reflects the concentration and drifting trend of the smoke itself. It can be seen that the present invention not only has a good reconstruction effect on traditional data sets, but also has good reconstruction quality for various actual video frames shot today, verifying its effectiveness compared to classic methods such as DCVC and its reliability in multi-scene applications.
[0089] according to Figure 8 The subjective results shown in the figure show that for the first group of examples, the white dotted lines and water marks on the road surface reconstructed by the comparison method DCVC at a low bit rate (<0.1bpp) are obviously blurred, while the white dotted line edges of the present invention are clear and continuous, and the water marks texture is consistent with the original. Figure 1 For the second set of examples, the road surface texture reconstructed by comparison methods such as DCVC at low bitrates (<0.1bpp) is overly smooth, with the edges of the white dashed lines blurred. However, the present invention preserves the road surface particle details intact, with the sharpness of the white dashed lines close to the original frame. This shows that the drone video reconstructed by the present invention at low bitrates retains complete details, which can meet the needs of practical scenarios such as video surveillance and nighttime inspections.
[0090] In summary, the present invention proposes a generative video coding method based on implicit inter-frame alignment, which is used to solve the problem of reduced perceptual quality of video data in a low bit rate environment.
[0091] The present invention utilizes an implicit inter-frame alignment strategy and effectively captures the potential features between frames through a diffusion model, thereby optimizing the processing and reconstruction capabilities of inter-frame information.
[0092] The present invention designs an adaptive spatiotemporal importance encoder, implements a more flexible compression strategy to adaptively adjust compression parameters, and further improves the reconstruction quality of the video.
[0093] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a specific implementation, and the modules or processes in the accompanying drawings are not necessarily necessary for implementing the present invention.
[0094] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:
[0095] An adaptive spatiotemporal importance encoder optimizes the quality of key regions while reducing the bit overhead of less important regions. An implicit inter-frame alignment strategy captures the underlying spatiotemporal relationships between frames, reducing the computational complexity of estimating explicit motion information and computational overhead. A perceptual loss function effectively improves the perceptual quality of video frames.
[0096] By combining a diffusion model with an implicit inter-frame alignment strategy and using decoded features as preconditions for pre-training the latent diffusion model, this method captures latent features between frames, thereby recovering target features and reconstructing video frames. This reduces the redundant information found in traditional motion estimation methods while better recovering detail and structural information between video frames. Furthermore, through an adaptive spatiotemporal importance encoder and a proposed spatiotemporal modulation unit, target features are scaled. A neural network is used to extract the importance weight of each frame, and the bitrate allocation is dynamically adjusted based on factors such as motion complexity and texture complexity, thereby improving reconstruction quality. Furthermore, by combining learning to perceive image block similarity, the method optimizes video visual quality in low-bitrate environments, overcoming the limitations of traditional evaluation methods.
[0097] From the above description of the embodiments, it is clear that those skilled in the art will clearly understand that the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the present invention.
[0098] The implementation methods in this specification are all described in a progressive manner, and the same or similar parts between the modules can be referred to each other, and each module focuses on the differences from other modules. In particular, for the device or system, since it is basically similar to the present invention, the description is relatively simple, and the relevant parts can be referred to the partial description of the present invention. The devices and systems described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0099] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A generative video coding method based on implicit inter-frame alignment, characterized in that: include: Step 1: The adaptive spatiotemporal importance encoder intelligently models the spatiotemporal characteristics of video frames and dynamically allocates bitrate. Step 2: The features of the previous frame are extracted through the implicit inter-frame alignment strategy and then fused with the information of the current frame. The information of the previous frame is transformed and aligned with the current frame through the structure of the U-Net network. Step 3: Combine perceptual image patch similarity constraints through perceptual loss function.
2. The method for generating video coding based on implicit inter-frame alignment according to claim 1, wherein: The step 1 includes: setting an original frame to be encoded, extracting the original frame through a feature extractor to obtain target features, and using reference frame features as prior information to encode the target features to obtain decoding features.
3. The generative video coding method based on implicit inter-frame alignment according to claim 2, wherein: The decoding feature is expressed as: where E(·) and D(·) represent the encoder and decoder with four-fold up- and down-sampling convolutional layers and two spatiotemporal modulation units, respectively. represents the reference feature; y i Represents target features; Decoding features.
4. The method for generating video coding based on implicit inter-frame alignment according to claim 1, wherein: The step 1 also includes: extracting spatiotemporal importance weights from reference features using a spatiotemporal attention network through a spatiotemporal modulation unit, which are used to scale current target features and perform bitrate allocation; wherein the reference features and target features are fused through a Concat operation, and then the spatiotemporal importance weights are obtained through a 3×3 convolutional layer with a ReLU activation function and a Sigmoid function.
5. The generative video coding method based on implicit inter-frame alignment according to claim 1, wherein: The step 1 further includes: using the spatiotemporal importance weight to balance the fusion feature f by the spatiotemporal modulation unit i and Conv(f i ) to extract the spatiotemporal features in the target features; the spatiotemporal importance weights and bit rate weights are further combined through the spatiotemporal modulation unit to perform dynamic bit rate allocation to obtain the scaled features.
6. The generative video coding method based on implicit inter-frame alignment according to claim 1, wherein: The step 2 includes: performing DDIM inverse diffusion noise addition on the reference feature to make it close to the distribution of Gaussian noise, and replacing the original random noise ε~N(0,1) with the initialization feature as the initial input of the diffusion model.
7. The generative video coding method based on implicit inter-frame alignment according to claim 1, wherein: The step 2 also includes: aligning the target features with the initialization features between frames and performing denoising optimization using U-Net to obtain high-quality latent space features; the state transfer process of the diffusion model It can be regarded as a process of gradually denoising from the initial features to the maximum likelihood features.
8. The method for generating video coding based on implicit inter-frame alignment according to claim 1, wherein: The maximum likelihood feature is expressed as: in, It represents the denoising prediction model based on U-Net, which is the total process of diffusion denoising; t represents the number of diffusion steps; represents the maximum likelihood feature; Represents the initialization feature.
9. The generative video coding method based on implicit inter-frame alignment according to claim 1, wherein: The step 3 includes: a loss function setting strategy, which is divided into two stages: the first stage uses the initial frame and a single forward frame to train the model to control the reconstruction loss from the initial frame to the forward frame; the second stage considers increasing the number of forward frames to learn the correlation between forward frames, further controlling the overall loss, and constructing a loss function based on the bit rate and reconstruction distortion results obtained from the two stages of training to obtain better compression effect.
10. The generative video coding method based on implicit inter-frame alignment according to claim 1, wherein: The step 3 comprises: During the state transfer process, the parameters of the U-Net network remain fixed, and the rate-distortion loss function Expressed as: Where λ represents the bit rate and the balance coefficient between reconstruction distortion, β represents the mean square error MSE loss and learning-aware patch similarity LPIPS loss The weight coefficient between represents the rate-distortion loss function; y i Represents target features; Represents mean square error MSE loss; represents the learning-perceptual image patch similarity LPIPS loss; x i (i∈{0,1,…n}) represents the input video sequence; Represents the reconstructed video sequence; under the above loss function setting strategy, The following formula will be obtained: in, To use the I frame and the previous frame for training, we get the bit rate, MSE loss, and LPIPS loss of the previous frame. To use the previous frame and the current frame for training, the bit rate, MSE loss and LPIPS loss of the current frame are obtained, y i represents the target feature; x i (i∈{0,1,…n}) represents the input video sequence; Represents the reconstructed video sequence.
Citation Information
Patent Citations
Compressed video quality enhancement method based on all-known network
CN115496683A
Fish behavior identification method and system based on behavior visual rhythm difference perception
CN118799961A
Cited By
Random window pixel alignment training method and system for video restoration
CN121437296A
A random window pixel alignment training method and system for video restoration
CN121437296B