Low-resolution video super-resolution enhancement method and system based on cognitive-diffusion model

By employing a cognitive-diffusion model-based video super-resolution approach, which combines data preprocessing, cognitive encoder training, and attention map registration, the image quality improvement and flicker artifact issues of low-resolution videos are addressed, achieving high-quality video super-resolution conversion applicable to various demanding scenarios.

CN120655505BActive Publication Date: 2025-12-09ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510446903.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-12-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Existing video super-resolution technologies suffer from insufficient detail recovery, blurring, and jagged edges when processing low-resolution videos. CNN methods have poor generalization ability in complex scenes, and diffusion models have difficulty guaranteeing temporal consistency and semantic accuracy in video processing, failing to meet the application requirements of high-demand scenarios.

Method used

A cognitive-diffusion model-based approach is adopted, which involves specific data preprocessing, cognitive encoder training, reference video generation, encoding and attention map registration, and latent representation decoding steps. By combining LDM and T2S models, inter-frame attention and temporal 3D residual blocks are introduced to optimize the decoding process and achieve high-quality video super-resolution conversion.

Benefits of technology

It solves the difficulties in improving the image quality of low-resolution videos and the problem of flickering artifacts. The generated videos have significant advantages in semantic accuracy and detail restoration, and are suitable for high-requirement scenarios such as high-definition video playback and professional film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655505B_ABST
    Figure CN120655505B_ABST
Patent Text Reader

Abstract

The application discloses a low-resolution video super-resolution enhancement method and system based on a cognitive-diffusion model. Through a series of method steps such as specific data preprocessing, cognitive encoder training, reference video generation, coding and attention map registration, and latent representation decoding, the transformation from a low-resolution video to a super-resolution video is realized, and the problems of low-resolution video quality improvement and flicker artifacts in the existing video processing process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and particularly relates to a low-resolution video super-resolution enhancement method and system based on cognitive coding and latent diffusion model. BACKGROUND

[0002] In the current digital era, digital video technology is developing at an unprecedented speed, and the application range of video is also expanding, widely penetrating into entertainment, education, medical treatment, security and many other fields, becoming a key medium for information dissemination and interaction. Under such circumstances, video super-resolution technology, as a core means to improve video quality, has experienced several important stages in its development.

[0003] In the early stage, video super-resolution technology mostly used interpolation algorithms to improve video resolution through pixel interpolation, which improved the viewing experience of low-resolution video to a certain extent. After the rise of machine learning technology, learning-based methods entered the field of video super-resolution, with CNN as a typical representative. It can automatically learn image features and master the mapping relationship between low and super-resolution images through a large amount of data training, making significant progress in restoring video details and improving resolution, and promoting the important development of video super-resolution technology. In recent years, deep learning technology has made major breakthroughs, and diffusion models have emerged. It can generate images or videos by learning data distribution and gradually denoising in the latent space, which can theoretically capture complex data features more accurately and bring new ideas and methods to video super-resolution.

[0004] Despite the above developments in video super-resolution technology, existing technologies still have many defects, and the present application is aimed at improving the following deficiencies.

[0005] (1) Traditional interpolation algorithms only estimate and fill in surrounding pixels, and cannot restore the lost detail information during resolution reduction, resulting in problems such as blurring and jaggedness in the generated video. In high-definition video playback, professional video production and other scenarios with high quality requirements, it is difficult to meet the needs and the quality improvement effect is limited.

[0006] (2) The method based on CNN has defects. On the one hand, its model structure makes it difficult to capture the global features of video in complex scenes, and its recovery ability is insufficient when facing videos with rich details and complex structures, resulting in poor detail performance. On the other hand, CNN is prone to overfitting, and its generalization ability is poor in scenes other than training data, which affects the stability of super-resolution processing and the reliability of actual application.

[0007] (3) The application of diffusion models faces challenges. First, there is the issue of temporal consistency. The randomness of the diffusion process makes it difficult to ensure the continuity of adjacent frames when processing videos. When increasing the resolution, the video is prone to flickering and stuttering, which is particularly prominent in scenarios with high requirements for smoothness, such as video entertainment and real-time monitoring. Second, there is the difficulty in utilizing semantic information. It cannot perform targeted super-resolution processing based on video semantics. The processed video may be semantically unreasonable and cannot meet the requirements of scenarios with high requirements for semantic accuracy, such as video surveillance analysis and medical image processing. This may lead to misjudgment of key information or affect the diagnosis of the disease. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a method and system for super-resolution enhancement of low-resolution videos based on a cognitive-diffusion model. Through a series of collaborative steps, including specific data preprocessing, cognitive encoder training, reference video generation, encoding and attention map registration, and latent representation decoding, the method achieves the conversion from low-resolution videos to super-resolution videos, solving the problems of difficult image quality enhancement of low-resolution videos and flicker artifacts in existing video processing.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A method for super-resolution enhancement of low-resolution videos based on a cognitive-diffusion model, characterized by the following steps:

[0011] S1. Obtain low-resolution video;

[0012] S2. Data preprocessing, the data preprocessing includes: S201. Video segmentation: segmenting the input low-resolution video, wherein for videos with high frame rates and rich dynamic changes, the sampling points are increased; for videos with low frame rates and relatively static videos, the sampling points are reduced; S202. Training data preparation; S203. Training the cognitive encoder;

[0013] S3. Reference Video Generation: The cognitive embedding generated by the cognitive encoder is used as input and fed into the LDM without introducing other parameters. This guides the diffusion process to generate a semantically consistent reference video, providing global semantic guidance for subsequent attention map registration.

[0014] S4. Reference Video Encoding and Attention Map Registration: This includes S401. Reference Video Encoding, which encodes the reference video using VAE-Encoder and maps the reference video to the latent space to obtain the latent representation; and S402. Attention Map Registration, which analyzes the latent representation of the encoded reference video, extracts attention information, uses the U-Net encoder to generate multi-scale control features, and associates the attention information of the reference video with the input video.

[0015] S5. Optimizing decoding and training to improve video quality.

[0016] Further, the S202. Training data preparation specifically comprises: processing ImageNet1000 data into 512x512 pixel size, then randomly selecting 2000 from the processed images to construct a test set, using an improved Real-ESRGAN algorithm to generate low-resolution versions of high-resolution images in the test set, and simultaneously using a BLIP2 multi-modal model to generate multi-modal captions for each high-resolution image; calculating the semantic similarity of images of the same class in the ImageNet1000 dataset based on the CLIP model, and using this similarity value as a supervision signal for the subsequent training of the reference image attention module.

[0017] Further, the S203. Training cognitive encoder specifically comprises:

[0018] Using the CLIP image encoder to represent the CLIP image embedding extracted from the low-resolution image as providing local structural information for super-resolution, where B, T i and C i represent batch size, label number and channel number, respectively;

[0019] At the same time, using the BLIP2 language encoder to extract the CLIP language embedding from the real annotation corresponding to the high-resolution HR image, and representing it as providing global semantic guidance for super-resolution images, where T l is the number of language labels, and C l is the language feature dimension; using a learnable query mechanism to compress the language embedding dimension: first define the number of learnable queries T e and ensure that T e <T l , and convert the original language embedding L to cognitive embedding using L ′ as the supervision signal.

[0020] The loss function for training the cognitive encoder is:

[0021]

[0022] where

[0023] where t cls represents the index position of the class label in the language embedding sequence, T e is a specific threshold for the number of learnable queries, and Padding() is a padding operation.

[0024] Further, the S3. reference video generation specific steps are as follows:

[0025] S301. Potential space representation: input the generated cognitive embedding into the LDM, which will map the cognitive embedding to a low-dimensional potential space through the encoder to extract key features; let the input cognitive embedding be z0, and the potential representation z obtained through the encoder; as a variant of DDPM, it aims to minimize the following objective function to predict artificial noise:

[0026]

[0027] Where ∈ is noise, z t is the potential representation at time t, and θ is the learnable parameter of the denoising network, θ includes convolution layer weights, time step embedding parameters, parameters of the conditional information fusion module, and weight matrix of the attention mechanism, representing the conditional information;

[0028] S302. Diffusion process: the diffusion process of the LDM gradually changes from the initial noise distribution to the distribution related to the cognitive embedding;

[0029] S303. T2S model construction and optimization: construct the T2S model for approximate inversion operation, which is based on the LDM architecture and inherits its image generation capability; expand the LDM to the video generation field through structural adjustment, and the model uses a 1x3x3 convolution kernel and a time attention mechanism, replacing self-attention with inter-frame attention; the inter-frame attention updates the current frame features according to the first frame and the current frame of the video through a specific projection matrix;

[0030] S304. Model initialization: initialize the W Q learnable projection matrix of the inter-frame attention and cross-attention, and additional time attention (fine-tune based on noise prediction of the input video;

[0031] S305. Video inversion based on shared unconditional embedding: use the fine-tuned T2S model for video inversion, and realize by optimizing the shared unconditional embedding is a globally shared learnable parameter used to uniformly control the semantic consistency of multiple frame generation, avoiding parameter redundancy caused by optimizing each frame independently.

[0032] Further, the VAE-Encoder uses a convolutional neural network architecture, which includes 5 convolutional layers with convolution kernel sizes of 3x3, 3x3, 5x5, 5x5, and 7x7, respectively, and strides of 1, 2, 1, 2, and 1, respectively; map the reference video to the potential space to obtain the potential representation, which has a shape of [32, 16, 32];

[0033] The U-Net encoder comprises a downsampling path and an upsampling path, the downsampling path is 4 convolutional layers, the upsampling path is 4 deconvolutional layers, the convolution kernel size is 3*3, and the step is 2; the attention information of the reference video is associated with the input video in a specific manner, including calculating an attention weight matrix, multiplying the attention weight matrix with the input video features, and focusing the model on the key part of the video.

[0034] Further, the S5. optimizing decoding and training to improve video quality comprises: S501. introducing a time 3D residual block: S502. adopting a training mode similar to a time U-Net, in the training process, keeping the pre-trained spatial layer parameters fixed, and focusing on targeted training of the newly added time layer. The S501. introducing a time 3D residual block comprises adding a time 3D residual block in the VAE-Decoder, and the residual block utilizes its unique three-dimensional structure to optimize the decoding process in depth. The S502. adopting a training mode similar to a time U-Net comprises that the time layer training adopts a hybrid loss function composed of an L1 loss, an LPIPS perceptual loss and an adversarial loss of a time PatchGAN discriminator.

[0035] Another object of the present application is to provide a low-resolution video super-resolution improvement system based on a cognitive-diffusion model, comprising a low-resolution video acquisition module, a data preprocessing module, a reference video generation module, a reference video encoding and attention map registration module, and an optimized decoding and training video quality improvement module, the low-resolution video super-resolution improvement system based on the cognitive-diffusion model is used to execute the above-mentioned low-resolution video super-resolution improvement method based on the cognitive-diffusion model.

[0036] Another object of the present application is to provide a computer readable storage medium storing one or more programs, the one or more programs causing a computer to execute the above-mentioned low-resolution video super-resolution improvement method based on the cognitive-diffusion model.

[0037] In combination with all the technical solutions, the present application has the following advantages compared with the prior art:

[0038] Through a series of method steps such as specific data preprocessing, cognitive encoder training, reference video generation, encoding and attention map registration, and latent representation decoding, the conversion from low-resolution video to super-resolution video is realized, and the problem of low-resolution video quality improvement and the problem of flicker artifacts in the existing video processing process are solved. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 A low-resolution video super-resolution improvement method based on a cognitive-diffusion model provided by the present application is shown in the flowchart;

[0040] Figure 2 A data preprocessing method flowchart is provided for the present application.

[0041] Figure 3 A reference video generation method flowchart is provided for the present application.

[0042] Figure 4 Another method flowchart for low-resolution video super-resolution enhancement based on a cognitive-diffusion model is provided. DETAILED DESCRIPTION

[0043] The embodiments will be further described below by way of examples and with reference to the accompanying drawings. Figures 1-4 It is obvious that the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0044] The present application provides a low-resolution video super-resolution enhancement method based on a cognitive-diffusion model, as shown in the following formula: Figures 1-4 The low-resolution video super-resolution enhancement method based on the cognitive-diffusion model includes the following steps:

[0045] S1. Obtain a low-resolution video.

[0046] Low-resolution video refers to a video with low resolution, usually lower than the common high-definition standard (such as 1080p or 1920x1080 pixels). As an example, a low-resolution movie segment is read from an early digital storage device. Due to the long time, its resolution is only 320x240, and the picture quality is affected by the storage conditions, with color deviation, scratches, etc.

[0047] S2. Data preprocessing:

[0048] S201. Video segmentation: segment the input low-resolution video. Cut the video into a series of picture frames at fixed time intervals or according to the key frame features of the video content. At the same time, according to the frame rate and content dynamic change characteristics of the video, sample each frame. For videos with high frame rate and rich dynamic changes, appropriately increase the sampling points; for videos with low frame rate and relatively static, reasonably reduce the sampling points to ensure sufficient coverage in the time dimension to capture the dynamic information of the video.

[0049] Preferably, according to the key frame features of the video content, the low-resolution video is segmented by using professional video analysis software. Since the movie scene changes frequently, key frames are selected at intervals of every 3-5 seconds for segmentation, while referring to the video frame rate (assuming 25 fps) and the dynamic change characteristics of the picture. In the action scene, the sampling points are appropriately increased, such as sampling once every 1-2 frames; in the relatively static dialogue scene, sampling once every 3-4 frames, to ensure that rich dynamic information can be captured.

[0050] S202. Training data preparation: preprocess the ImageNet1000 dataset, first unify the image to 512x512 pixel size, then randomly select 2000 from the processed image to build a test set. Use the improved Real-ESRGAN algorithm (a blind super-resolution model trained based on pure synthetic data, generate diversified degradation modes through CycleGAN (Cycle-Consistent Generative Adversarial Network), combined with perceptual loss and adversarial training to improve generalization ability), then degrade the high-resolution images in the test set to generate low-resolution (LR) versions, and use BLIP2 multi-modal model to generate multi-modal captions for each high-resolution (HR) image. Calculate the semantic similarity of images of the same class in the ImageNet1000 dataset based on the CLIP model, and use this similarity value as a supervision signal for the subsequent reference image attention module training. The ImageNet1000 dataset is derived from the ILSVRC2012 competition, containing 1000 image classes, about 1000 sample images per class, widely used as a benchmark dataset in the field of computer vision for model training and performance evaluation.

[0051] S203. Train the cognitive encoder: use the CLIP image encoder to represent the CLIP image embedding extracted from the low-resolution (LR) image as I∈R B×Ti×Ci , providing local structural information for super-resolution, ensuring that the restored image is visually consistent with the low-resolution (LR) image input, where B, T i and C i represent batch size, label number and channel number respectively. In this embodiment, the specific parameter settings are B=32, T i =77, C i =768. At the same time, use the BLIP2 language encoder to extract the CLIP language embedding from the real annotation corresponding to the high-resolution (HR) image, and represent it as L∈R B×Tl×Cl after processing by the BLIP2 model, providing global semantic guidance for super-resolution images, avoiding semantic misunderstanding due to low-resolution (LR) image blur, where T lFor the number of language tokens, C l For the language feature dimension. In the cognitive encoder, we use a learnable query mechanism to compress the dimension of the language embedding. The specific operation is as follows: first, define the number of learnable queries T e (T e = 5) and ensure that T e <T l , the CLIP language embedding L is converted into cognitive embedding E through linear projection E B×Te×Cl , the supervision signal adopts L ′ , the specific rule is Where t cls represents the index position of the category token in the language embedding sequence, T e The specific threshold of the number of learnable queries is used to compress the dimension of the language embedding, and Padding() is the padding operation. When the number of tokens available for supervision is less than T e , the category token L[t cls ] is used for padding to ensure that the length of the supervision signal reaches T e . Through this supervision method, the cognitive encoder can learn a more comprehensive semantic representation than a single category. The invention suggests using the T e tokens before the category token L[t cls ] for supervision, where T e is set to 5 because these tokens retain all previous information and provide richer and more comprehensive semantic guidance for training. If the number of supervision tokens is insufficient, the category token is used for padding to ensure the integrity of the supervision information. The loss function used to train the cognitive encoder uses mean square error (MSE) as follows: Where is the cognitive embedding, L ′ is the supervision signal, is the square of the L2 norm

[0052] The initial learning rate is set to 0.001, and the Adam optimizer is used, with the learning rate decaying to 0.9 every 5 epochs.

[0053] S3. Reference video generation: the cognitive embedding generated by the cognitive encoder is input into the LDM (Latent Diffusion Model) without introducing other parameters, guiding the diffusion process to generate a semantic consistent reference video, providing global semantic guidance for subsequent attention map registration. The specific steps are as follows:

[0054] S301. Latent space representation: input the generated cognitive embedding into the LDM (Latent Diffusion Model), guide the diffusion process through the cognitive embedding, constrain the semantic consistency between video frames, and solve the flicker artifact problem caused by the randomness of traditional diffusion models. LDM will map the cognitive embedding to a low-dimensional latent space through the LDM encoder to extract key features. As a variant of DDPM, LDM optimizes the model by predicting artificial noise and learns the distribution of the data.

[0055] Preferably, the LDM maps the cognitive embedding to a low-dimensional latent space through the latent diffusion model encoder. The shape of the input cognitive embedding is [32, 77, 768], and the latent representation obtained through the latent diffusion model encoder becomes [32, 32, 64] to capture the key features of the input data. Let the input cognitive embedding be z0, and the latent representation obtained through the latent diffusion model encoder be z. As a variant of DDPM, it aims to minimize the following objective function to predict artificial noise:

[0056]

[0057] where ∈ is noise, z t is the latent representation at time t, and θ is the learnable parameter of the denoising network (including convolutional layer weights, time step embedding parameters, parameters of the conditional information fusion module, and attention mechanism weight matrix), represents the conditional information. During training, 1000 iterations are performed, and the batch size is 16 for each iteration.

[0058] S302. Diffusion process:

[0059] The diffusion process of LDM gradually changes from the initial noise distribution to the distribution related to the cognitive embedding. At each step of the diffusion, the latent representation is updated according to a specific formula, and the parameter related to the diffusion step controls the amount of noise added at each step. When processing real images, the image is first encoded into the latent space, then denoised and decoded back into the image space. The reverse process of DDIM sampling also has a corresponding formula for backward derivation of the latent representation.

[0060] Preferably, the diffusion process of LDM is a Markov chain that moves from the initial noise distribution to the distribution related to the cognitive embedding. At diffusion step t, the latent representation z t is updated according to the following formula:

[0061]

[0062] where α t is related to the diffusion step and determines the amount of noise added, z t is the latent representation at time t, and ∈ is the noise, θ is the learnable parameter of the denoising network. The real image is first encoded into the latent space, and then decoded back to the image space after denoising. The reverse denoising process adopts the DDIM sampling strategy, and the denoising network with the learning parameter θ The original latent representation is gradually recovered. The specific reverse update formula is:

[0063]

[0064] where α t Related to the diffusion step, determines the amount of noise added, z t is the latent representation at time t, ∈ is the noise, θ is the learnable parameter of the denoising network.

[0065] The diffusion process uses a guided noise prediction network ∈ θ Learn the noise distribution under semantic constraints, effectively solve the semantic deviation problem of traditional diffusion models in video processing.

[0066] S303. T2S (Text to Set) model construction and optimization: A T2S (Text to Set) model is constructed for approximate reverse operation. The model is based on the latent diffusion model (LDM) architecture, inheriting its image generation capability. Through structural adjustment, LDM is extended to the field of video generation, and T2S model uses 1x3x3 mode convolution kernel and time attention mechanism, replacing self-attention (Self Attention) with inter-frame attention (Frame Attention). Frame attention updates the current frame features according to the first frame and the current frame of the video through a specific projection matrix:

[0067] Q = W Q v i , K = W K v0, V = W V v0

[0068] where W Q , W K , W V is a learnable projection matrix, v i is the current frame feature, v0 is the first frame feature of the video, Q is the Query matrix, used to calculate the attention weight between the current frame and the first frame, guiding the model to pay attention to the key semantic information in the first frame, K is the Key matrix, as a semantic anchor of the reference frame, providing a reference for attention weight calculation, V is the Value matrix, carrying the core semantic information of the first frame, which is weighted and participates in the update of the current frame feature after attention weight.

[0069] When constructing the model, the parameters of the attention mechanism should be optimized to ensure the quality and semantic consistency of the generated video.

[0070] initialized as a random normal distribution with a standard deviation of 0.01 and optimized through backpropagation during the training process. When constructing the T2S model, the mean square error (MSE) loss function is used to optimize the attention mechanism parameters, and the learning rate is set to 0.0001. The model is predicted through frame-by-frame processing and multiple calculations. When constructing, the attention mechanism needs to be optimized to ensure video quality and semantic consistency.

[0071] S304. T2S model initialization:

[0072] The W Q and additional temporal attention are fine-tuned based on the noise prediction of the input video. After initialization and fine-tuning, the T2S model can generate a set of semantically consistent images while maintaining frame quality, achieving approximate inversion.

[0073] Model expansion helps inter-frame semantic consistency, but it will damage the quality of the T2S model generated, as the self-attention parameter calculates the frame correlation without pre-training. Therefore, the W Q and additional temporal attention are fine-tuned, and during the fine-tuning process, other layer parameters are fixed, only the attention mechanism related parameters are adjusted, fine-tuned for 50 epochs, the learning rate is set to 0.00001, and the noise prediction of the input video is used. After initialization and fine-tuning, the T2S model can generate a set of semantically consistent images while maintaining frame quality, achieving approximate inversion.

[0074] S305. Video inversion based on shared unconditional embedding:

[0075] The fine-tuned T2S model is used for video inversion, and the shared unconditional embedding is optimized to achieve is a globally shared learnable parameter that is used to uniformly control the semantic consistency of multiple frames, avoiding parameter redundancy caused by independent optimization of each frame. In the inversion process, the latent feature covers the dimension channel of the video frame number, and the DDIM inversion is obtained The optimization goal of the unconditional embedding is:

[0076]

[0077] Here is the expected latent feature, is the T2S model generated feature. The inter-frame attention uses two latent features to calculate the next step feature, which is shared by all frames, and finally the complete reference video is obtained.

[0078] S4. Reference video encoding and attention map registration

[0079] S401. Reference video encoding

[0080] After generating the reference video, to better associate the input video key information, the reference video encoding and attention map registration are performed. The reference video is encoded by the VAE-Encoder to map the reference video to the latent space and obtain the latent representation. This latent representation compresses the video information, facilitating subsequent processing. The VAE-Encoder selects a convolutional neural network architecture, which includes 5 convolutional layers with kernel sizes of 3x3, 3x3, 5x5, 5x5, and 7x7, respectively, and step sizes of 1, 2, 1, 2, and 1, respectively. The reference video is mapped to the latent space to obtain the latent representation, which has a shape of [32, 16, 32].

[0081] S402, attention map registration

[0082] The reference decoupling guidance strategy determines the attention information extraction method. The encoded reference video latent representation is analyzed to extract attention information. Through a specific way, a multi-scale control feature is generated using a U-Net encoder (Unified Neural Network Encoder) to associate the reference video attention information with the input video, allowing the T2S model to focus on the key parts of the video.

[0083] The U-Net encoder includes a downsampling path (4 convolutional layers) and an upsampling path (4 deconvolutional layers), both with a kernel size of 3x3 and a step size of 2. The reference video attention information is associated with the input video through a specific way, such as calculating an attention weight matrix, which is multiplied with the input video features to make the model focus on the key parts of the video.

[0084] S5. Optimizing decoding and training to improve video quality

[0085] The latent space obtained from the previous processing is restored using the optimized decoder to generate HR (High Resolution) images. This decoder accurately maps the information in the latent space back to the high-resolution image space, reasonably reconstructing the information in spatial and temporal dimensions to ensure that the generated video frames have rich details and accurate structures.

[0086] S501. Introducing a time 3D residual block:

[0087] The fine-tuned decoder is used to restore the latent space to high-resolution images. Since the VAE-Decoder within the LDM framework is only trained on images, it is prone to flickering artifacts when decoding latent sequences. To this end, a temporal 3D residual block is added to the VAE-Decoder, which utilizes its unique three-dimensional structure to optimize the decoding process. It not only captures the change information of video frames in the temporal dimension, but also strengthens the correlation between adjacent frames, dynamically adjusts the decoding output, and enhances the consistency of low-level features. The attention map registers the key region information extracted, guiding the temporal 3D residual block to prioritize the temporal coherence of these regions during decoding, avoiding excessive smoothing in non-key regions. These residual blocks can optimize the decoding process, reduce flickering, and improve the visual stability of the video.

[0088] S502. To further improve video quality, this step adopts a training mode similar to temporal U-Net. During training, the pre-trained spatial layer parameters are kept fixed, and the newly added temporal layer is focused on targeted training.

[0089] The temporal layer training adopts a hybrid loss function composed of L1 loss, LPIPS perceptual loss, and adversarial loss of temporal PatchGAN discriminator. Among them, the L1 loss calculates the difference between the generated image and the original image at the pixel level, prompting the generated image to be as close as possible to the original image in terms of pixel value, ensuring that the basic structure and content of the image are accurate and accurate. LPIPS perceptual loss simulates the evaluation criteria of the human visual system for image quality from the professional perspective of human eye visual perception. The adversarial loss of the temporal PatchGAN discriminator learns through mutual confrontation with the generator, aiming to improve the authenticity and temporal consistency of the generated image.

[0090] By continuously adjusting the weights of the three losses in the hybrid loss function and optimizing the T2S model parameters, the performance of the decoder is comprehensively optimized, and finally high-quality super-resolution videos are generated.

[0091] Preferably, the L1 loss weight is set to 0.5, the LPIPS perceptual loss weight is set to 0.3, and the temporal PatchGAN discriminator adversarial loss weight is set to 0.2. By adjusting these losses, the performance of the decoder is optimized. During training, each epoch contains 100 batches, each batch contains 8 video segments, and each video segment contains 16 frames. After 50 epochs of training, high-quality super-resolution videos are generated. After processing, the original 320x240 resolution movie segments are upgraded to 1080p resolution.

[0092] The application solves the problem of low-resolution video super-resolution by organically fusing cognitive-semantic constraint mechanism, latent diffusion guidance generation, space-time attention alignment and dynamic time modeling. Specifically, first, a cognitive encoder is constructed to fuse CLIP image embedding and language embedding to generate compact semantic representation, and to learn deep semantic association beyond a single category through dynamic padding supervision strategy. The module adopts batch size B = 32 (balances training efficiency and memory occupation), label number T i = 77 (retains complete spatial semantic information) and learnable query number T e = 5 (compresses language embedding dimension and improves effective information amount of supervision signal), to provide semantic anchor points for subsequent diffusion process. Then, a T2S model architecture is proposed to replace the self-attention of LDM with inter-frame attention mechanism and design shared unconditional embedding to realize multi-frame semantic consistency. The module breaks through the limitation of traditional diffusion model independent processing of frames by latent space dimension [32, 32, 64] (reduces computational complexity) and inter-frame attention projection matrix initialization standard deviation 0.01 (improves convergence speed). At the same time, a U-Net encoder is developed to generate multi-scale control features [32, 16, 32], to realize attention map registration of reference video and input video, to capture cross-frame semantic dependence combined with 3×3×1 convolution kernel (reduces computational amount) and time attention module, to form semantic perception-space positioning double closed loop. Finally, 3D residual block is introduced to construct space-time feature fusion path, and hybrid loss function is adopted to balance pixel-level, perception-level and time consistency constraints. Each technical means constitutes an inseparable cooperative system: the cognitive encoder provides semantic prior for the diffusion process, the inter-frame attention mechanism is the basis for realizing shared embedding, the attention map registration provides spatial reference for time modeling, and the hybrid loss function forms a generation-discrimination closed loop optimization. The application has significant advantages in semantic accuracy and detail recovery through deep cooperation of cognitive embedding and diffusion generation, realizes video semantic understanding through self-supervised cognitive encoder, and is more suitable for text-free annotation scene.

[0093] Correspondingly, the present application provides a low-resolution video super-resolution enhancement system based on a cognitive-diffusion model, comprising a low-resolution video acquisition module, a data preprocessing module, a reference video generation module, a reference video encoding and attention map registration module, and an optimized decoding and training video quality enhancement module. The low-resolution video super-resolution enhancement system based on the cognitive-diffusion model is used to execute the low-resolution video super-resolution enhancement method based on the cognitive-diffusion model described above. It should be noted that those skilled in the art should understand that the implementation functions of each module shown in the embodiment of the low-resolution video super-resolution enhancement system based on the cognitive-diffusion model can be understood with reference to the related description of the low-resolution video super-resolution enhancement method based on the cognitive-diffusion model. The functions of each module shown in the embodiment of the low-resolution video super-resolution enhancement system based on the cognitive-diffusion model can be realized by a program (executable instructions) running on a processor, or by a specific logic circuit.

[0094] Correspondingly, the present application also provides a computer-readable storage medium, wherein computer executable instructions are stored, and the computer executable instructions are executed by a processor to implement the method embodiments of the present application. The computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0095] In addition, it should be understood that the above description is only the preferred embodiment of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the protection scope of one or more embodiments of the present application.

Claims

1. A method for low-resolution video super-resolution enhancement based on a cognitive-diffusion model, characterized in that, Comprising the following steps: S1. Obtain a low-resolution video; S2. Data preprocessing, comprising: S201. Video segmentation: segmenting the input low-resolution video, wherein for videos with high frame rate and rich dynamic changes, the sampling points are increased; for videos with low frame rate and relatively static, the sampling points are reduced; S202. Training data preparation; S203. Training cognitive encoder; S3. Reference video generation: input the cognitive embedding generated by the cognitive encoder into the LDM without introducing other parameters, and guide the diffusion process to generate a semantically consistent reference video to provide global semantic guidance for subsequent attention map registration; S4. Reference video encoding and attention map registration: comprising S401. Reference video encoding, encoding the reference video with a VAE-Encoder to map the reference video to a latent space and obtain a latent representation; and S402. Attention map registration, analyzing the encoded reference video latent representation, extracting attention information, and generating multi-scale control features using a U-Net encoder to associate the attention information of the reference video with the input video; S5. Optimizing decoding and training to improve video quality; Wherein, the specific steps of S3. Reference video generation are as follows: S301. Latent space representation: input the generated cognitive embedding into the LDM, which will map the cognitive embedding to a low-dimensional latent space through the encoder to extract key features; let the input cognitive embedding be z0, and the latent representation z obtained after encoding; as a variant of DDPM, it aims to minimize the following objective function to predict artificial noise: where ∈ is noise, z t is the latent representation at time t, and θ are the learnable parameters of the denoising network, including the convolutional layer weights, the time step embedding parameters, the parameters of the conditional information fusion module, and the weight matrix of the attention mechanism, denotes the conditional information; S302. Diffusion process: the diffusion process of LDM is to gradually change from the initial noise distribution to the distribution related to the cognitive embedding; S303. T2S model construction and optimization: a T2S model is constructed for approximate inverse operation, which is based on the LDM architecture and inherits its image generation capability; the LDM is expanded to the video generation field through structural adjustment, and the model uses a 1x3x3 convolution kernel and a temporal attention mechanism to replace self-attention with inter-frame attention; the inter-frame attention updates the current frame features through a specific projection matrix based on the first frame and the current frame of the video; S304. Model initialization: W for inter-frame attention and cross-attention Q learnable projection matrix, and additional temporal attention are fine-tuned based on noise prediction of the input video; S305. Video inversion based on shared unconditional embedding: video inversion using the fine-tuned T2S model, by optimizing the shared unconditional embedding Implementation, is a globally shared learnable parameter that is used to uniformly control the semantic consistency of multi-frame generation, avoiding parameter redundancy caused by independent optimization of each frame.

2. The low-resolution video super-resolution enhancement method based on the cognitive-diffusion model according to claim 1, wherein, The S202. Training data preparation specifically comprises: processing the ImageNet1000 data into 512x512 pixel size, then randomly selecting 2000 images from the processed images to construct a test set, using an improved Real-ESRGAN algorithm to degrade the high-resolution images in the test set to generate low-resolution versions, and simultaneously using a BLIP2 multi-modal model to generate multi-modal captions for each high-resolution image; based on the CLIP model, the semantic similarity of images of the same type in the ImageNet1000 data set is calculated, and this similarity value is used as a supervision signal for the subsequent training of the reference image attention module.

3. The method of claim 1, wherein the method is based on a cognitive- diffusion model. The S203. Training cognitive encoder specifically comprises: Embedding the CLIP image extracted from the low resolution image into a representation using a CLIP image encoder is provides local structure information for super-resolution, where B, T i and C i denote batch size, label number, and channel number, respectively; Meanwhile, the CLIP language embedding is extracted from the real label corresponding to the high-resolution HR image using the BLIP2 language encoder, and is expressed as The global semantic guidance is provided for the super-resolution image, wherein T l is the number of language labels, C l is the language feature dimension; the language embedding dimension is compressed by using a learnable query mechanism: first, define the number of learnable queries T e and ensure that T e <T l , and the original language embedding L is converted into a cognitive embedding The supervision signal adopts L ′ , The loss function for training the cognitive encoder is: wherein where t cls denotes the index position of the category label in the language embedding sequence, T e is a specific threshold of the number of learnable queries, and Padding() is a padding operation.

4. The low-resolution video super-resolution enhancement method based on the cognitive-diffusion model of claim 1, wherein, The VAE-Encoder uses a convolutional neural network architecture, including 5 convolutional layers, with convolution kernel sizes of 3x3, 3x3, 5x5, 5x5, and 7x7, respectively, and step sizes of 1, 2, 1, 2, and 1, respectively; maps the reference video to the latent space to obtain a latent representation, which has a shape of [32, 16, 32]; The U-Net encoder includes a downsampling path and an upsampling path, the downsampling path is 4 convolutional layers, and the upsampling path is 4 deconvolutional layers, with a convolution kernel size of 3x3 and a step size of 2; the attention information of the reference video is associated with the input video in a specific way, including calculating an attention weight matrix and multiplying it with the input video features to make the model focus on the key parts of the video.

5. The method of claim 1, wherein the method is based on a cognitive- diffusion model. The S5. optimizing decoding and training to improve video quality includes: S501. introducing a time 3D residual block: S502. using a time U-Net-like training mode, during the training process, keeping the pre-trained spatial layer parameters fixed, and focusing on targeted training of the newly added time layer.

6. The low-resolution video super-resolution enhancement method based on cognitive-diffusion model according to claim 5, wherein, The S501. introducing a time 3D residual block includes adding a time 3D residual block in the VAE-Decoder, which uses its unique three-dimensional structure to optimize the decoding process in depth.

7. The method of claim 5, wherein the method is based on a cognitive- diffusion model. The S502. using a time U-Net-like training mode includes using a hybrid loss function composed of L1 loss, LPIPS perceptual loss, and time PatchGAN discriminator adversarial loss for time layer training.

8. A low-resolution video super-resolution boosting system based on a cognitive-spreading model, characterized in that, The low-resolution video super-resolution enhancement system based on the cognitive-diffusion model includes a low-resolution video acquisition module, a data preprocessing module, a reference video generation module, a reference video encoding and attention map registration module, and an optimization decoding and training video quality improvement module, and is used to execute the method of any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, One or more programs are stored, which cause the computer to execute the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Video super-resolution reconstruction method and system based on multi-scale local self-attention

    CN115082308A

  • Method and apparatus for generating image for video super resolution

    KR102702246B1