Video quality enhancement method, apparatus, device, and storage medium

By using a diffusion model and a local perception cross-attention module, high-quality videos are generated, which solves the problem of insufficient fine-grained detail generation in existing technologies and achieves improved video quality and enhanced robustness.

CN120640028BActive Publication Date: 2026-01-27BEIJING ZHIXIANG FUTURE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510730118.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2026-01-27
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing video enhancement technologies have limitations in generating novel, fine-grained details, making it difficult to produce high-quality, richly detailed videos.

Method used

The video keyframes are compressed into low-quality latent codes using a diffusion model to generate high-quality latent codes. High-quality videos are then generated through noise enhancement, optical flow models, and mask enhancement techniques. Finally, texture alignment is performed using a local perception cross-attention module to generate high-quality videos.

Benefits of technology

The generated video not only contains rich fine-grained details, but also maintains the coherence between frames, improving the detail generation capability and robustness of the video image quality, and can stably output high-quality results under different input conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640028B_ABST
    Figure CN120640028B_ABST
Patent Text Reader

Abstract

The application provides a video quality enhancement method, device, equipment and storage medium, the method comprising: obtaining an initial video of low quality; obtaining a key frame of the initial video; compressing the key frame into a low-quality hidden code through a diffusion model, obtaining noise hidden codes of each time step according to the low-quality hidden code, and generating a high-quality hidden code according to the noise hidden codes of each time step; enhancing the high-quality hidden code and the low-quality hidden code to obtain an enhanced high-quality hidden code and an enhanced low-quality hidden code; and generating a high-quality video according to the enhanced high-quality hidden code and the enhanced low-quality hidden code. After obtaining the initial video of low quality, the method of the application obtains the low-quality hidden code through the key frame of the initial video, and then generates the high-quality hidden code based on the noise hidden codes of each time step obtained from the low-quality hidden code. After enhancing the high-quality hidden code and the low-quality hidden code, the high-quality video is generated, which not only contains fine-grained details, but also needs to maintain the coherence of inter-frame details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus, device, and storage medium for enhancing video quality. Background Technology

[0002] Video enhancement technology aims to improve video quality by analyzing and processing video pixel information.

[0003] Existing video quality enhancement technologies mainly use sliding window methods or recursive methods to aggregate information from low-resolution video frames using convolutional neural networks or Transformers to predict high-resolution frames.

[0004] While such methods excel at amplifying original visual content, they have limitations in generating novel, fine-grained details. Summary of the Invention

[0005] To address one of the aforementioned technical deficiencies, this application provides a video quality enhancement method, apparatus, device, and storage medium.

[0006] A first aspect of this application provides a video quality enhancement method, the method comprising:

[0007] Obtain low-quality initial video;

[0008] Obtain keyframes from the initial video;

[0009] The keyframes are compressed into low-quality hidden codes using a diffusion model. Noise hidden codes for each time step are obtained based on the low-quality hidden codes. High-quality hidden codes are then generated based on the noise hidden codes for each time step.

[0010] Enhance the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code.

[0011] High-quality video is generated based on enhanced high-quality steganography and enhanced low-quality steganography.

[0012] Optionally, the noise hidden code at any time step t

[0013] Where, α t and σ t is a hyperparameter, and ∈ represents Gaussian noise.

[0014] Optionally, a high-quality steganography is generated based on the noise steganography at each time step, including:

[0015] High-quality hidden codes are generated through iterative denoising at multiple time steps.

[0016] Among them, through the formula Perform denoising at any time step t;

[0017] z t To calculate the hidden code of the image after denoising at time step t using a diffusion model, For the noise code at time step t, γ t For denoising weights.

[0018] Optionally,

[0019] Where T is the maximum time step.

[0020] Optionally, the high-quality hidden code and the low-quality hidden code are enhanced to obtain enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0021] Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to both high-quality and low-quality implicit codes, resulting in noise-enhanced high-quality and enhanced low-quality implicit codes.

[0022] Based on the enhanced low-quality latent code, the high-quality latent code after noise enhancement is distorted by an optical flow model to generate a distorted high-quality latent code that is aligned with the motion of the initial video.

[0023] Masking enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code.

[0024] Optionally, a high-quality video is generated based on the enhanced high-quality steganography and the enhanced low-quality steganography, including:

[0025] Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively.

[0026] Based on the query features, key features, and value features, a texture-aligned feature embedding vector Conv(AV)+Q is generated; where A is the attention weight. Q represents the query feature, K represents the key feature, v represents the value feature, B represents the local attention bias matrix, and Norm(·) represents the feature normalization function. is the scaling factor, T is the transpose operator, and softmax(·) is the normalization function;

[0027] Generate high-quality videos based on feature embedding vectors.

[0028] Optionally,

[0029] in, b ij represents the elements of the local attention bias matrix, where i is the row identifier, j is the column identifier, and d is the length of the local neighborhood corresponding to the local attention.

[0030] A second aspect of this application provides a video quality enhancement device, the device comprising:

[0031] The first acquisition module is used to acquire low-quality initial video;

[0032] The second acquisition module is used to acquire keyframes of the initial video;

[0033] The first generation module is used to compress keyframes into low-quality hidden codes using a diffusion model, obtain noisy hidden codes for each time step based on the low-quality hidden codes, and generate high-quality hidden codes based on the noisy hidden codes for each time step.

[0034] The enhancement module is used to enhance high-quality and low-quality steganography to obtain enhanced high-quality and enhanced low-quality steganography.

[0035] The second generation module is used to generate high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code.

[0036] A third aspect of this application provides an electronic device, comprising:

[0037] Memory;

[0038] Processor; and

[0039] Computer programs;

[0040] The computer program is stored in the memory and configured to be executed by the processor to implement the method described in the first aspect above.

[0041] In a fourth aspect, this application provides a computer-readable storage medium having a computer program stored thereon; the computer program is executed by a processor to implement the method described in the first aspect above.

[0042] This application provides a video quality enhancement method, apparatus, device, and storage medium. The method includes: acquiring a low-quality initial video; acquiring keyframes of the initial video; compressing the keyframes into a low-quality hidden code using a diffusion model; obtaining noise hidden codes at each time step based on the low-quality hidden code; generating a high-quality hidden code based on the noise hidden codes at each time step; enhancing the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code; and generating a high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code. The method of this application, after acquiring a low-quality initial video, obtains a low-quality hidden code from the keyframes of the initial video, then generates a high-quality hidden code based on the noise hidden codes at each time step obtained from the low-quality hidden code, and after enhancing the high-quality hidden code and the low-quality hidden code, generates a high-quality video. This high-quality video not only contains fine-grained details but also maintains the coherence of details between frames. Attached Figure Description

[0043] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0044] Figure 1 A flowchart illustrating a video quality enhancement method provided in an embodiment of this application;

[0045] Figure 2 A schematic diagram of a video quality enhancement model provided in an embodiment of this application;

[0046] Figure 3 A schematic diagram of a local sensing cross-attention module provided in an embodiment of this application;

[0047] Figure 4 This is a schematic diagram of the structure of a video quality enhancement device provided in an embodiment of this application;

[0048] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0049] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0050] In developing this application, the inventors discovered that video quality enhancement technology aims to improve video quality by analyzing and processing video pixel information. Existing video quality enhancement technologies primarily utilize sliding window methods or recursive methods, employing convolutional neural networks or Transformers to aggregate information from low-resolution video frames to predict high-resolution frames. While these methods excel at amplifying original visual content, they have limitations in generating novel, fine-grained details.

[0051] To address the aforementioned problems, this application provides a video quality enhancement method, apparatus, device, and storage medium. The method includes: acquiring a low-quality initial video; acquiring keyframes of the initial video; compressing the keyframes into low-quality hidden codes using a diffusion model; obtaining noise hidden codes at each time step based on the low-quality hidden codes; generating high-quality hidden codes based on the noise hidden codes at each time step; enhancing the high-quality hidden codes and low-quality hidden codes to obtain enhanced high-quality hidden codes and enhanced low-quality hidden codes; and generating a high-quality video based on the enhanced high-quality hidden codes and enhanced low-quality hidden codes. The method of this application, after acquiring a low-quality initial video, obtains low-quality hidden codes from the keyframes of the initial video, then generates high-quality hidden codes based on the noise hidden codes at each time step obtained from the low-quality hidden codes. After enhancing the high-quality hidden codes and low-quality hidden codes, a high-quality video is generated. This high-quality video not only contains fine-grained details but also maintains the coherence of details between frames.

[0052] See Figure 1 This embodiment provides a video quality enhancement method, which can enhance video quality through... Figure 2 The video quality enhancement model shown is implemented as follows:

[0053] 101, Get a low-quality initial video.

[0054] In step 101, a low-quality initial video will be acquired, such as... Figure 2 LQ Video (Low-quality Video) refers to the initial video of low quality.

[0055] 102. Obtain the keyframes of the initial video.

[0056] In step 102, the following will be obtained: Figure 2 Keyframes of the initial LQ Video.

[0057] For example, upsampling the initial LQ Video (e.g.) Figure 2 In the Upsampling process, the first frame after each upsampling is determined as the key frame. Then, in step 102, the first frame after LQ Video upsampling will be obtained.

[0058] 103. The keyframes are compressed into low-quality hidden codes using a diffusion model. Noise hidden codes for each time step are obtained based on the low-quality hidden codes. High-quality hidden codes are generated based on the noise hidden codes for each time step.

[0059] In step 103, the keyframes are compressed into low-quality implicit codes z using a diffusion model (such as a variational autoencoder based on the diffusion model), and the noise implicit codes for each time step are obtained based on the low-quality implicit codes z. Based on the noise code at each time step Generate high-quality hidden code y.

[0060] The high-quality implicit code y is a semantically aligned high-quality reference image implicit code y.

[0061] The process of compressing keyframes into low-quality hidden codes can be achieved through a variational autoencoder of a diffusion model. For example, a variational autoencoder of a diffusion model can be used to compress video keyframes into low-quality hidden codes z.

[0062] The noise code at each time step is obtained from the low-quality hidden code z. Based on the noise code at each time step The process of generating a high-quality hidden code y can be achieved through reverse diffusion using a diffusion model.

[0063] For example, at each time step t (t = T, T-1, ..., 1, where T is the preset maximum time step) during denoising, the image hidden code z after denoising at time step t is obtained through back-diffusion of the diffusion model. t Image hidden code z t With a noise code having a positive noise-adding time step t Perform a weighted summation to utilize the hidden codes from the noise. Using semantic components to modulate image hidden codes z t High-quality hidden code y is generated through iterative denoising over multiple time steps (i.e., T time steps).

[0064] Therefore, in the specific implementation, the noise hidden code at each time step is obtained based on the low-quality hidden code z. At any time step t, the noise hidden code

[0065] Where, α t and σ t is a hyperparameter, such as the constant hyperparameter of the noise scheduling function in DDPM (Denoising Diffusion Probabilistic Models). ∈ represents Gaussian noise, such as standard Gaussian noise.

[0066] When generating a high-quality hidden code based on the noise hidden code at each time step, high-quality hidden code y can be generated by iterative denoising at multiple time steps.

[0067] Among them, through the formula Perform denoising at any time step t.

[0068] z t To calculate the hidden code of the image after denoising at time step t using a diffusion model, The noise code for time step t.

[0069] γ t For denoising weights, γ t According to the definition of the cosine function of t, such as

[0070] Where T is the maximum time step, which can be the maximum time step used when pre-training the diffusion model.

[0071] Since the maximum value of the denoising time step t is the maximum time step T of the pre-trained diffusion model (e.g., T = 1000), according to γ t The value monotonically decreases from 0.5 to 0 as t decreases from its maximum value T to 0, ensuring that more semantic information from the original frame is introduced in the early stages of the denoising stage, while adding as much low-level texture detail as possible in the later stages. Finally, through iterative denoising, a high-quality reference image implicit code y with semantic alignment and rich texture is generated.

[0072] Step 103 can be achieved through Figure 2 The Semantics-Aligning Image Diffuser implementation in [the example is missing]. For instance, in step 103, the Semantics-Aligning Image Diffuser first compresses the video keyframes into a hidden code (i.e., a low-quality hidden code z) using a variational autoencoder of a diffusion model. This low-quality hidden code z is as follows: Figure 2 In the Z, i is the keyframe identifier, and z is the keyframe identifier. i Let Z be the low-quality latent code corresponding to the i-th keyframe, and Z be the set of low-quality latent codes corresponding to all keyframes (i = 1, 2, ..., I, where I is the total number of keyframes). Gaussian noise is added to the latent code (i.e., the low-quality latent code z). Then, a reverse diffusion is performed using an image diffusion model, and the noise latent code (i.e., the noise latent code) is reduced by weighted summation. ) and denoising hidden code (i.e. z) t By combining these, a high-quality implicit code y (i.e., ...) is generated. Figure 2 Image Reference (in the image reference).

[0073] It should be noted that the hidden code mentioned in this embodiment and subsequent real-time examples refers to the features obtained by encoding the image using a trained variational autoencoder, which can be used to decode the image and reconstruct it. All inputs Figure 2The video and images in the video quality enhancement model shown are processed into implicit code form by a variational autoencoder.

[0074] 104. Enhance the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code.

[0075] In step 104, the high-quality hidden code y and the low-quality hidden code Z are enhanced to obtain the enhanced high-quality hidden code Y and the enhanced low-quality hidden code Z.

[0076] In step 104, noise enhancement, temporal motion enhancement, and mask enhancement are performed. The implementation process of step 104 is as follows:

[0077] 1. Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to both the high-quality and low-quality hidden codes, resulting in noise-enhanced high-quality and enhanced low-quality hidden codes.

[0078] For example, by performing forward diffusion using a diffusion model, n steps of Gaussian noise are added to both the high-quality implicit code y and the low-quality implicit code z, resulting in the enhanced high-quality implicit code y″. n And the enhanced low-quality hidden code Z n The enhanced low-quality hidden code Z n That is, the enhanced low-quality hidden code Z.

[0079] This section is for noise enhancement, which can be performed on both high-quality hidden code y and low-quality hidden code z to simulate video enhancement scenarios under different noise levels.

[0080] For example, by performing forward diffusion using a diffusion model, n steps of Gaussian noise are added to both the high-quality implicit code y and the low-quality implicit code z to generate the enhanced high-quality implicit code y″. n And the enhanced low-quality hidden code Z n .

[0081] In practice, visual fidelity and quality can be balanced by adjusting the noise level. For example, the noise level can vary randomly.

[0082] 2. Based on the enhanced low-quality hidden code, the high-quality hidden code after noise enhancement is distorted by the optical flow model to generate a distorted high-quality hidden code that is aligned with the motion of the initial video.

[0083] For example, based on the enhanced low-quality hidden code (i.e., the enhanced low-quality hidden code Z). n The enhanced high-quality hidden code y″ is obtained through optical flow models (such as the RAFT model). n Perform a warp transformation to generate a high-quality hidden code y′ that is aligned with the initial video motion. n .

[0084] This section focuses on temporal motion enhancement, which improves the temporal coherence of the video. Optical flow estimation techniques are used to perform motion compensation on low-quality videos, and motion information is then incorporated into the diffusion process of high-quality videos.

[0085] For example, the optical flow of low-quality video is estimated using the RAFT optical flow model, and the optical flow is used to enhance the high-quality hidden code y″. n Perform a distortion transformation based on the enhanced low-quality implicit code (i.e., the enhanced low-quality implicit code Z). n Generate a high-quality hidden code y′ after warp transformation aligned with the motion of low-quality video. n .

[0086] It's important to note that the RAFT model is a pre-trained optical flow extraction model; simply inputting the video into the RAFT model will extract the optical flow. The warp transformation can align the content between video frames. Therefore, a high-quality reference image implicit code sequence can be obtained by transforming the optical flow extracted from different frames into a high-quality reference image implicit code sequence, which can then be aligned with the motion of the low-quality video.

[0087] 3. Perform mask enhancement on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code.

[0088] For example, the high-quality hidden code y′ after the warp transformation n Perform mask enhancement to generate the mask-enhanced hidden code Y. n The hidden code Y after the mask enhancement n That is, the enhanced high-quality hidden code Y.

[0089] This section describes mask enhancement, which can improve the model training effect through mask operations and avoid the video quality enhancement method and model provided in this embodiment from relying too much on a single reference image.

[0090] For example, the high-quality hidden code y′ after the warp transformation n A masking operation is randomly applied to set some pixel blocks to zero (e.g., randomly selecting a certain proportion (e.g., 0.5) of pixel blocks and setting them to 0), generating a masked high-quality reference image implicit code (i.e., the enhanced high-quality implicit code Y). The enhanced high-quality implicit code Y preserves the complete high-quality reference image implicit code sequence, ensuring that the model can utilize the complete reference image information.

[0091] Step 104 can be achieved Figure 2 The Augmentor implementation in the code yields an enhanced, high-quality implicit code Y (i.e., Figure 2 Y in n , Figure 2 middle Y is the enhanced high-quality hidden code corresponding to the i-th keyframe. nEnhanced high-quality latent code for all keyframes) and enhanced low-quality latent code Z (i.e. Figure 2 Z in n , Figure 2 middle Z is the enhanced low-quality hidden code corresponding to the i-th keyframe. n (The enhanced low-quality occult code set corresponding to all keyframes).

[0092] 105. Generate high-quality video based on enhanced high-quality hidden codes and enhanced low-quality hidden codes.

[0093] In step 105, a high-quality video is generated based on the enhanced high-quality hidden code Y and the enhanced low-quality hidden code Z.

[0094] The implementation process of step 105 is as follows:

[0095] 1. Extract feature blocks from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and use them as query features, key features and value features, respectively.

[0096] Feature blocks can be extracted from the enhanced high-quality implicit code Y and the enhanced low-quality implicit code Z, and used as query feature Q, key feature K, and value feature V, respectively. For example, three features can be obtained using three independent mapping networks, which are Q, K, and V, respectively. That is, Q, K, and V have the same source but are mapped to different spaces.

[0097] 2. Based on the query features, key features, and value features, generate a texture-aligned feature embedding vector e as Conv(AV)+Q.

[0098] Where A is the attention weight. Q represents the query feature, K represents the key feature, V represents the value feature, B represents the local attention bias matrix, and Norm(·) represents the feature normalization function. is the scaling factor, T is the transpose operator, and softmax(·) is the normalization function.

[0099] in,

[0100] b ij represents the elements of the local attention bias matrix, where i is the row identifier, j is the column identifier, and d is the length of the local neighborhood corresponding to the local attention.

[0101] The process of generating the texture-aligned feature embedding vector e described above can be achieved through... Figure 3The local perceptual cross-attention module is shown in the diagram. This module enables texture alignment between a high-quality reference image and a low-quality video. Specifically, feature blocks are extracted from the enhanced high-quality implicit code Y and the enhanced low-quality implicit code Z, serving as the query feature Q, key feature K, and value feature V, respectively. During attention computation, a local attention bias matrix B is added to the attention map to limit the attention range to the local neighborhood d, thereby enhancing the alignment effect of local texture features.

[0102] 3. Generate high-quality videos based on feature embedding vectors.

[0103] For example, a high-quality video can be generated based on the feature embedding vector e.

[0104] In practical implementation, after obtaining the feature embedding vector e, it can be done through... Figure 2 The Texture-AligningVideo Controller adds it to the output of subsequent encoder blocks for multi-scale vision-guided learning.

[0105] Specifically, the Texture-Aligning Video Controller uses the texture-aligned feature embedding vector e output by the local perception cross-attention module, and obtains visual guidance features through feature propagation of the cascaded encoder block. The visual guidance features are then used to adjust the denoising process of the video diffusion model to generate high-quality video.

[0106] In step 105, a texture-aligned feature embedding vector e is generated through cross-attention calculation and added to the output of subsequent encoder blocks for multi-scale visual guidance learning. Finally, the Texture-AligningVideo Controller uses the texture-aligned feature embedding vector output by the local perception cross-attention module to learn visual guidance features through cascaded encoder blocks, thereby adjusting the denoising process of the video diffusion model and ultimately generating a high-quality video.

[0107] The video quality enhancement method provided in this embodiment can greatly improve the ability to generate fine-grained details in high-quality videos. By introducing a high-quality reference image (i.e., an enhanced high-quality implicit code Y) and using a local perceptual cross-attention module for texture alignment, it can generate richer fine-grained details, such as textures and edges. This makes the generated video more visually realistic and delicate, solving the shortcomings of existing technologies in detail generation.

[0108] The video quality enhancement method provided in this embodiment has strong robustness. By using noise enhancement, temporal motion enhancement, and masking enhancement, it can simulate various complex input scenarios, thereby enhancing the robustness of the model. This enables the video quality enhancement method and model provided in this embodiment to output high-quality video results more stably when facing different input conditions and degradation scenarios, solving the problem of insufficient robustness in existing technologies.

[0109] The video quality enhancement method provided in this embodiment allows for flexible control over the generation quality and fidelity. By adjusting the noise level parameter, a flexible balance can be struck between video quality and original visual fidelity. This enables video enhancement technology to better meet diverse needs.

[0110] This embodiment provides a video quality enhancement method, which involves: acquiring a low-quality initial video; acquiring keyframes from the initial video; compressing the keyframes into low-quality hidden codes using a diffusion model; obtaining noise hidden codes at each time step based on the low-quality hidden codes; generating high-quality hidden codes based on the noise hidden codes at each time step; enhancing both the high-quality and low-quality hidden codes to obtain enhanced high-quality and enhanced low-quality hidden codes; and generating a high-quality video based on the enhanced high-quality and enhanced low-quality hidden codes. In this embodiment, after acquiring the low-quality initial video, the method obtains low-quality hidden codes from the keyframes of the initial video, then generates high-quality hidden codes based on the noise hidden codes at each time step obtained from the low-quality hidden codes. After enhancing both the high-quality and low-quality hidden codes, a high-quality video is generated. This high-quality video not only contains fine-grained details but also maintains the coherence of details between frames.

[0111] Based on the same inventive concept as video quality enhancement methods, this embodiment provides a video quality enhancement device, such as... Figure 4 As shown, the device includes:

[0112] The first acquisition module 401 is used to acquire low-quality initial video.

[0113] The second acquisition module 402 is used to acquire keyframes of the initial video.

[0114] The first generation module 403 is used to compress keyframes into low-quality hidden codes through a diffusion model, obtain noise hidden codes at each time step based on the low-quality hidden codes, and generate high-quality hidden codes based on the noise hidden codes at each time step.

[0115] Enhancement module 404 is used to enhance high-quality steganography and low-quality steganography to obtain enhanced high-quality steganography and enhanced low-quality steganography.

[0116] The second generation module 405 is used to generate a high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code.

[0117] Wherein, the noise hidden code at any time step t

[0118] Where, α t and σ t is a hyperparameter, and ∈ represents Gaussian noise.

[0119] The generation of high-quality hidden codes based on the noise hidden codes at each time step includes:

[0120] High-quality hidden codes are generated through iterative denoising at multiple time steps.

[0121] Among them, through the formula Perform denoising at any time step t.

[0122] z t To calculate the hidden code of the image after denoising at time step t using a diffusion model, For the noise code at time step t, γ t For denoising weights.

[0123] in,

[0124] Where T is the maximum time step.

[0125] Among them, the high-quality hidden codes and low-quality hidden codes are enhanced to obtain enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0126] Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to both high-quality and low-quality implicit codes, resulting in noise-enhanced high-quality and enhanced low-quality implicit codes.

[0127] Based on the enhanced low-quality latent code, a distortion transformation is performed on the noise-enhanced high-quality latent code using an optical flow model to generate a distortion-transformed high-quality latent code that is aligned with the initial video motion.

[0128] Masking enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code.

[0129] Among these, high-quality video is generated based on enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0130] Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively.

[0131] Based on the query features, key features, and value features, a texture-aligned feature embedding vector Conv(AV)+Q is generated. Here, A represents the attention weight. Q represents the query feature, K represents the key feature, V represents the value feature, B represents the local attention bias matrix, and Norm(·) represents the feature normalization function. is the scaling factor, T is the transpose operator, and softmax(·) is the normalization function.

[0132] Generate high-quality videos based on feature embedding vectors.

[0133] in,

[0134] in, b ij represents the elements of the local attention bias matrix, where i is the row identifier, j is the column identifier, and d is the length of the local neighborhood corresponding to the local attention.

[0135] The device provided in this embodiment acquires a low-quality initial video, obtains a low-quality hidden code through the key frames of the initial video, and then generates a high-quality hidden code based on the noise hidden codes at each time step obtained from the low-quality hidden code. After enhancing the high-quality hidden code and the low-quality hidden code, a high-quality video is generated. This high-quality video not only contains fine-grained details, but also needs to maintain the coherence of details between frames.

[0136] Based on the same inventive concept as video quality enhancement methods, this embodiment provides an electronic device, which is as follows: Figure 5 As shown, it includes: memory 501, processor 502, and computer program.

[0137] The computer program is stored in memory 501 and configured to be executed by processor 502 to implement the above-described video quality enhancement method.

[0138] Specifically,

[0139] Get a low-quality initial video.

[0140] Obtain keyframes from the initial video.

[0141] The keyframes are compressed into low-quality hidden codes using a diffusion model. Noise hidden codes for each time step are obtained from the low-quality hidden codes, and high-quality hidden codes are generated from the noise hidden codes for each time step.

[0142] Enhanced high-quality and low-quality hidden codes are obtained by enhancing them.

[0143] High-quality video is generated based on enhanced high-quality steganography and enhanced low-quality steganography.

[0144] Wherein, the noise hidden code at any time step t

[0145] Where, αt and σ t is a hyperparameter, and ∈ represents Gaussian noise.

[0146] The generation of high-quality hidden codes based on the noise hidden codes at each time step includes:

[0147] High-quality hidden codes are generated through iterative denoising at multiple time steps.

[0148] Among them, through the formula Perform denoising at any time step t.

[0149] z t To calculate the hidden code of the image after denoising at time step t using a diffusion model, For the noise code at time step t, γ t For denoising weights.

[0150] in,

[0151] Where T is the maximum time step.

[0152] Among them, the high-quality hidden codes and low-quality hidden codes are enhanced to obtain enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0153] Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to both high-quality and low-quality implicit codes, resulting in noise-enhanced high-quality and enhanced low-quality implicit codes.

[0154] Based on the enhanced low-quality latent code, a distortion transformation is performed on the noise-enhanced high-quality latent code using an optical flow model to generate a distortion-transformed high-quality latent code that is aligned with the initial video motion.

[0155] Masking enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code.

[0156] Among these, high-quality video is generated based on enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0157] Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively.

[0158] Based on the query features, key features, and value features, a texture-aligned feature embedding vector Conv(AV)+Q is generated. Here, A represents the attention weight. Q represents the query feature, K represents the key feature, V represents the value feature, B represents the local attention bias matrix, and Norm(·) represents the feature normalization function. is the scaling factor, T is the transpose operator, and softmax(·) is the normalization function.

[0159] Generate high-quality videos based on feature embedding vectors.

[0160] in,

[0161] in, b ij represents the elements of the local attention bias matrix, where i is the row identifier, j is the column identifier, and d is the length of the local neighborhood corresponding to the local attention.

[0162] The electronic device provided in this embodiment has a computer program executed by a processor to obtain a low-quality initial video. The low-quality hidden code is obtained through the key frames of the initial video. Then, a high-quality hidden code is generated based on the noise hidden code of each time step obtained from the low-quality hidden code. After enhancing the high-quality hidden code and the low-quality hidden code, a high-quality video is generated. This high-quality video not only contains fine-grained details, but also needs to maintain the coherence of details between frames.

[0163] Based on the same inventive concept as the video quality enhancement method, this embodiment provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the aforementioned video quality enhancement method.

[0164] Specifically,

[0165] Get a low-quality initial video.

[0166] Obtain keyframes from the initial video.

[0167] The keyframes are compressed into low-quality hidden codes using a diffusion model. Noise hidden codes for each time step are obtained from the low-quality hidden codes, and high-quality hidden codes are generated from the noise hidden codes for each time step.

[0168] Enhanced high-quality and low-quality hidden codes are obtained by enhancing them.

[0169] High-quality video is generated based on enhanced high-quality steganography and enhanced low-quality steganography.

[0170] Wherein, the noise hidden code at any time step t

[0171] Where, α t and σ t is a hyperparameter, and ∈ represents Gaussian noise.

[0172] The generation of high-quality hidden codes based on the noise hidden codes at each time step includes:

[0173] High-quality hidden codes are generated through iterative denoising at multiple time steps.

[0174] Among them, through the formula Perform denoising at any time step t.

[0175] z t To calculate the hidden code of the image after denoising at time step t using a diffusion model, For the noise code at time step t, γ t For denoising weights.

[0176] in,

[0177] Where T is the maximum time step.

[0178] Among them, the high-quality hidden codes and low-quality hidden codes are enhanced to obtain enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0179] Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to both high-quality and low-quality implicit codes, resulting in noise-enhanced high-quality and enhanced low-quality implicit codes.

[0180] Based on the enhanced low-quality latent code, a distortion transformation is performed on the noise-enhanced high-quality latent code using an optical flow model to generate a distortion-transformed high-quality latent code that is aligned with the initial video motion.

[0181] Masking enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code.

[0182] Among these, high-quality video is generated based on enhanced high-quality hidden codes and enhanced low-quality hidden codes, including:

[0183] Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively.

[0184] Based on the query features, key features, and value features, a texture-aligned feature embedding vector Conv(AV)+Q is generated. Here, A represents the attention weight. Q represents the query feature, K represents the key feature, V represents the value feature, B represents the local attention bias matrix, and Norm(·) represents the feature normalization function. is the scaling factor, T is the transpose operator, and softmax(·) is the normalization function.

[0185] Generate high-quality videos based on feature embedding vectors.

[0186] in,

[0187] in, b ij represents the elements of the local attention bias matrix, where i is the row identifier, j is the column identifier, and d is the length of the local neighborhood corresponding to the local attention.

[0188] The computer-readable storage medium provided in this embodiment allows a computer program thereon to be executed by a processor to obtain a low-quality initial video. After obtaining the low-quality hidden code through the key frames of the initial video, a high-quality hidden code is generated based on the noise hidden code at each time step obtained from the low-quality hidden code. After enhancing the high-quality hidden code and the low-quality hidden code, a high-quality video is generated. This high-quality video not only contains fine-grained details but also needs to maintain the coherence of details between frames.

[0189] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0190] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0191] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0192] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0193] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0194] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for enhancing video quality, characterized in that, The method includes: Obtain low-quality initial video; Obtain the keyframes of the initial video; The keyframes are compressed into low-quality hidden codes using a diffusion model. Noise hidden codes for each time step are obtained based on the low-quality hidden codes. High-quality hidden codes are then generated based on the noise hidden codes for each time step. The high-quality hidden code and the low-quality hidden code are enhanced to obtain enhanced high-quality hidden code and enhanced low-quality hidden code. Generate high-quality video based on enhanced high-quality steganography and enhanced low-quality steganography; The enhancement of the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code includes: Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to the high-quality hidden code and the low-quality hidden code respectively, resulting in noise-enhanced high-quality hidden code and enhanced low-quality hidden code. Based on the enhanced low-quality latent code, the high-quality latent code after noise enhancement is distorted by an optical flow model to generate a distorted high-quality latent code that is aligned with the motion of the initial video. Mask enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code; The step of generating a high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code includes: Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively. Based on the query features, key features, and value features, a texture-aligned feature embedding vector is generated. ;in, For attention weights, , To query features, Key features, For value characteristics, This is the local attention bias matrix. For the characteristic normalization function, Scaling factor It is the transpose operator. This is the normalization function; High-quality videos are generated based on the feature embedding vectors.

2. The method according to claim 1, characterized in that, Noise code at any time step t ; in, and For hyperparameters, It is Gaussian noise.

3. The method according to claim 1, characterized in that, The generation of high-quality hidden codes based on the noise hidden codes at each time step includes: High-quality hidden codes are generated through iterative denoising at multiple time steps. Among them, through the formula Perform denoising at any time step t; To calculate the time step using the diffusion model Hidden code in the denoised image For time step Noise-based hidden codes For denoising weights.

4. The method according to claim 3, characterized in that, ; in, This represents the maximum time step.

5. The method according to claim 1, characterized in that, ; in, , These are elements of the local attention bias matrix. For row identifiers, For column identifiers, This represents the length of the local neighborhood corresponding to the local attention.

6. A video quality enhancement device, characterized in that, The device includes: The first acquisition module is used to acquire low-quality initial video; The second acquisition module is used to acquire keyframes of the initial video; The first generation module is used to compress the keyframe into a low-quality hidden code using a diffusion model, obtain the noise hidden code for each time step based on the low-quality hidden code, and generate a high-quality hidden code based on the noise hidden code for each time step. An enhancement module is used to enhance the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code. The second generation module is used to generate high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code; The enhancement of the high-quality hidden code and the low-quality hidden code to obtain enhanced high-quality hidden code and enhanced low-quality hidden code includes: Forward diffusion is performed using a diffusion model to add multi-step Gaussian noise to the high-quality hidden code and the low-quality hidden code respectively, resulting in noise-enhanced high-quality hidden code and enhanced low-quality hidden code. Based on the enhanced low-quality latent code, the high-quality latent code after noise enhancement is distorted by an optical flow model to generate a distorted high-quality latent code that is aligned with the motion of the initial video. Mask enhancement is performed on the high-quality hidden code after the distortion transformation to generate enhanced high-quality hidden code; The step of generating a high-quality video based on the enhanced high-quality hidden code and the enhanced low-quality hidden code includes: Feature blocks are extracted from the enhanced high-quality implicit code and the enhanced low-quality implicit code, and used as query features, key features and value features, respectively. Based on the query features, key features, and value features, a texture-aligned feature embedding vector is generated. ;in, For attention weights, , To query features, Key features, For value characteristics, This is the local attention bias matrix. For the characteristic normalization function, Scaling factor It is the transpose operator. This is the normalization function; High-quality videos are generated based on the feature embedding vectors.

7. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, It stores a computer program thereon; the computer program is executed by a processor to implement the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Video data processing method and device, electronic equipment and readable storage medium

    CN118283297A

  • Diffusion model reversal image reconstruction method and device based on domain self-adaption

    CN118887138A