Ultrahigh-definition image and video real-time enhancement method and system and medium

By integrating video frame extraction and multimodal information on 4K videos, the gallery is enhanced by using visual language models and color retrieval, the problem of dynamic scene changes in high-resolution videos is solved, and the efficient enhancement of videos and optimization of computing resources is achieved.

CN119991530AActive Publication Date: 2025-05-13SUN YAT SEN UNIV +1

Patent Information

Application Number
CN202510204852.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing 4K video enhancement technology is difficult to effectively deal with changes in dynamic scenes in high-resolution videos, resulting in inconsistent timing, flickering or color drifting, and the computing resources are consumed, making it difficult to process a large amount of detailed information.

Method used

By extracting video frames on the source video, combining visual language model and color retrieval enhancement gallery, the text features and color features of the image are obtained, and input them into the DA2Net model to enhance and fusion of multimodal information to generate recovered and enhanced images.

Benefits of technology

It realizes the accuracy enhancement of each frame when processing 4K video, ensuring the overall visual fluency and color consistency of the video, reducing the consumption of computing resources, and improving the naturalness and coherence of the enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991530A_ABST
    Figure CN119991530A_ABST
Patent Text Reader

Abstract

The invention provides an ultra-high-definition image and video real-time enhancement method and system and a medium. The ultra-high-definition image and video real-time enhancement method comprises the following steps: performing video frame extraction operation on a source video to obtain a corresponding image sequence; inputting the image sequence into a visual language model, and performing image text description on the image sequence through the visual language model to obtain a text feature corresponding to each image; constructing a color retrieval enhancement gallery, and inputting the image sequence into the color retrieval enhancement gallery for retrieval to obtain color features needing to be supplemented; inputting the image sequence, the text features and the color features into a DA2Net model to obtain an output result of the DA2Net model; and performing convolution operation and image sequence fusion on an output result of the DA2Net model to obtain a recovered image, and generating an enhanced video based on the recovered image. According to the method, image enhancement processing of the ultra-high-definition video can be efficiently carried out on the premise that color features and image details are reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ultra-high-definition video enhancement technology, and in particular to a method, system and medium for real-time enhancement of ultra-high-definition images and videos. Background Art

[0002] Creating images with high contrast, vivid colors, and rich details is a core goal of photography. However, uneven lighting conditions in the scene often lead to over- or underexposure, resulting in loss of details and color distortion. This degradation in image quality may also have an adverse impact on downstream tasks such as object detection and image segmentation. To address this challenge, current methods usually reconstruct abnormally exposed areas by inferring details from surrounding pixels. However, the details and color information of such reconstructions often lack authenticity and are prone to hallucination effects, which in turn distort the representation of real scenes.

[0003] Classic color correction methods in image enhancement include histogram equalization, gamma correction, and Retinex algorithm. Traditional methods extend histogram equalization from grayscale images to color images to obtain better saturation, while gamma correction methods use nonlinear transformations to enhance dark details and suppress bright areas. Retinex-based methods enhance details and colors by estimating reflectivity and illumination. However, these methods often introduce noise or cause color distortion under complex lighting or uneven brightness.

[0004] Currently, CSEC uses a UNet-based network structure combined with a pseudo-normal feature generator to extract color features, and corrects the color shift of overexposed and underexposed areas through color shift estimation (COSE) and color modulation (COMO) modules. By utilizing a customized cross-attention mechanism and an optimized loss function, CSEC is able to enhance the brightness and color accuracy of images under suboptimal lighting conditions. However, its performance depends heavily on the quality of the pseudo-normal feature map, and the training process is relatively complex.

[0005] Deep learning-based methods, such as CSEC, have been widely used in various image enhancement tasks, including high dynamic range (HDR) reconstruction, low-light enhancement, and underwater image enhancement. These methods perform well under challenging conditions such as extreme lighting or poor visibility. However, since different tasks require targeted designs, there is still a lack of a unified framework to address various image enhancement tasks.

[0006] With the popularity of 4K video (resolution of 3840×2160), the enhancement of high-resolution video faces higher technical requirements. Existing 4K video enhancement technologies mainly rely on feature extraction in the spatial and temporal domains, but they usually have the following shortcomings in the process of high-resolution video processing: First, traditional enhancement methods fail to effectively cope with dynamic scene changes in high-resolution videos, which easily leads to inconsistent timing, flickering or color drift. Second, 4K videos usually contain a lot of detailed information, which places extremely high demands on computing power. Traditional methods are prone to computing bottlenecks when processing large-scale data. Finally, existing enhancement methods often ignore the semantic information in the video content, especially in complex lighting changes or scene transitions, which easily leads to unnatural or incoherent enhancement effects. Summary of the invention

[0007] The purpose of the present invention is to overcome the above-mentioned technical deficiencies and provide a method, system and medium for real-time enhancement of ultra-high-definition images and videos, which can ensure that each frame can be accurately enhanced through effective fusion of frame information and enhancement of multimodal information when processing ultra-high-definition videos, while ensuring the overall visual smoothness and color consistency of the video.

[0008] In order to achieve the above technical objectives, in a first aspect, the technical solution of the present invention provides a method for real-time enhancement of ultra-high-definition images and videos, the method for real-time enhancement of ultra-high-definition images and videos comprises the following steps:

[0009] Performing a video frame extraction operation on the source video to obtain a corresponding image sequence;

[0010] Inputting the image sequence into a visual language model, and performing image text description on the image sequence through the visual language model to obtain text features corresponding to each image;

[0011] Constructing a color retrieval enhancement gallery, inputting the image sequence into the color retrieval enhancement gallery to retrieve and obtain the color features that need to be supplemented;

[0012] The image sequence, the text features, and the color features are input into the D 2 Net model, and obtain the DA 2 The output of Net model;

[0013] The DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0014] Compared with the prior art, the beneficial effects of the present invention include:

[0015] First, the 4K video is frame-extracted to obtain a certain number of video frames. Then, given a color-distorted video frame, the video frame is simultaneously input into the Color Retrieval Enhanced Gallery (CRAG) and the open-source visual language model CogVLM. The Color Retrieval Enhanced Gallery (CRAG) retrieves the most needed k word vectors to supplement the semantic information of the input image in the color semantic space. The CogVLM visual language model extracts text prior information from the image. Then, the original image, the retrieved color information, and the text prior information are jointly input into the DAR. 2 In the Net model, DA 2 The output of the Net model is added to the original input image to generate a restored and enhanced image; finally, the enhanced video frames are reassembled to restore the video. Since video content usually has high dynamic variability and complex spatiotemporal information, this method can ensure that each frame can be accurately enhanced when processing 4K video through effective fusion of frame information and enhancement of multimodal information, while ensuring the overall visual smoothness and color consistency of the video.

[0016] According to some embodiments of the present invention, the image sequence is input into a visual language model, and the image sequence is described by the visual language model to obtain the text features corresponding to each image, including the steps of:

[0017] inputting the image sequence into a visual language model;

[0018] Input a prompt, and the visual language model performs a text description on the image sequence based on the prompt;

[0019] The text description of the image sequence is compiled by a text compiler to obtain text features.

[0020] According to some embodiments of the present invention, constructing a color retrieval enhanced gallery comprises the steps of:

[0021] Applying a Gaussian blur kernel to the image sequence to blur texture details and preserve overall color distribution;

[0022] Extracting image feature vectors from the image sequence using an embedding model ResNet;

[0023] The image feature vector is stored in a vector database to obtain the color retrieval enhanced gallery.

[0024] According to some embodiments of the present invention, the image sequence is input into the color retrieval enhancement gallery to retrieve and obtain the color features that need to be supplemented, including the steps of:

[0025] Given a feature vector of the image for query, construct a classifier to query the index of the most desired color vector;

[0026] The most needed k word vectors are selected, and based on the image feature vector and the classifier, the color retrieval enhancement gallery is searched to obtain the color features that need to be supplemented.

[0027] According to some embodiments of the present invention, the DA 2 Net models include:

[0028] Multiple residual Transformer attention blocks and convolutional layers. The residual Transformer attention blocks are used to extract and fuse features.

[0029] The residual Transformer attention block includes multiple local Transformer attention layers and convolutional layers;

[0030] The local Transformer attention layer includes: a differential proxy attention mechanism, layer normalization and a multi-layer perceptron, and the differential proxy attention mechanism is used to calculate attention features.

[0031] According to some embodiments of the present invention, the differential proxy attention mechanism introduces a set of additional proxy tokens A into the traditional attention module, represented as a four-tuple (Q, A, K, V). The proxy token A first acts as a proxy for the query token Q, responsible for aggregating information from K and V, and then broadcasting it back to Q, represented as:

[0032] O A =Attn S (Q,A,Attn S (A,K,V))=σ(QA T )σ(AK T )V

[0033] in, are the query, key, and value matrices, is the proxy token and σ(·) represents the Softmax function.

[0034] According to some embodiments of the present invention, extracting and fusing features in a residual Transformer attention block comprises the steps of:

[0035] Given the input feature F of the i-th residual Transformer attention block i,0 , first use L local Transformer attention layers (LTAL) to extract the intermediate features F i,1 ,F i,2 ,…,F i,L , the process is expressed as:

[0036] F i,j =LTAL i,j (Fi,j-1 ),j=1,2,…,L

[0037] Among them, LTAL i,j (·) represents the jth local Transformer attention layer in the i-th residual Transformer attention block. Then, a convolutional layer is added before the residual connection. The output of the residual Transformer attention block is expressed as:

[0038] F i,out =Conv i (F i,L )+F i,0

[0039] Among them, Conv i (·) is the convolutional layer in the i-th residual Transformer attention block.

[0040] According to some embodiments of the present invention, the DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, including the steps of:

[0041] The restored image is obtained using the following formula:

[0042]

[0043] in, is the final restored image, F f It's DA 2 The output of the Net model, X in is the image sequence, and Conv(·) is the convolution operation.

[0044] In a second aspect, the technical solution of the present invention provides an ultra-high-definition image and video real-time enhancement system, comprising:

[0045] A video frame extraction module is used to perform a video frame extraction operation on the source video to obtain a corresponding image sequence;

[0046] A text feature extraction module, which is in communication with the video frame extraction module and is used to input the image sequence into a visual language model, and to perform image text description on the image sequence through the visual language model to obtain text features corresponding to each image;

[0047] A color feature extraction module is connected to the video frame extraction module for constructing a color retrieval enhancement gallery, and inputs the image sequence into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented;

[0048] DA 2Net model, inputting the image sequence, the text features, and the color features into the D 2 Net model, and obtain the DA 2 The output of Net model;

[0049] Video enhancement module, with the DA 2 Net model communication connection, the DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0050] In a third aspect, the technical solution of the present invention provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the ultra-high-definition image and video real-time enhancement method as described in any one of the first aspects.

[0051] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, wherein the abstract drawings are identical to one of the drawings in the specification:

[0053] Figure 1 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention;

[0054] Figure 2 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention;

[0055] Figure 3 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention;

[0056] Figure 4 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention;

[0057] Figure 5 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention;

[0058] Figure 6 A method for real-time enhancement of ultra-high-definition images and videos provided by an embodiment of the present invention 2 Net model structure diagram. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0060] It should be noted that, although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0061] Reference Figure 1 , Figure 2 and Figure 6 , Figure 1 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention; Figure 2 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by one embodiment of the present invention; Figure 6 A method for real-time enhancement of ultra-high-definition images and videos provided by an embodiment of the present invention 2 Net model structure diagram. The ultra-high-definition image and video real-time enhancement method includes but is not limited to the following steps, and the numbers of the following steps should not be understood as having a sequence:

[0062] Step S110, performing a video frame extraction operation on the source video to obtain a corresponding image sequence;

[0063] Step S120, inputting the image sequence into the visual language model, and performing image text description on the image sequence through the visual language model to obtain the corresponding text features of each image;

[0064] Step S130, constructing a color retrieval enhancement gallery, inputting the image sequence into the color retrieval enhancement gallery to retrieve and obtain the color features that need to be supplemented;

[0065] Step S140: input the image sequence, text features, and color features into the D 2 Net model, and obtain DA 2 The output of Net model;

[0066] Step S150: 2 The output of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0067] In some embodiments, the method for real-time enhancement of ultra-high-definition images and videos includes the following steps: performing a video frame extraction operation on a source video to obtain a corresponding image sequence; inputting the image sequence into a visual language model, and performing an image text description on the image sequence through the visual language model to obtain the corresponding text features of each image; constructing a color retrieval enhancement gallery, and inputting the image sequence into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented; inputting the image sequence, text features, and color features into the D 2 Net model, and obtain DA 2 Net model output; 2 The output of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0068] Step S110, extract frames from the input 4K video to obtain a corresponding image sequence, receive an input 4K resolution video, and perform frame extraction on the input 4K video. Specifically, extract each frame of the video at a predetermined frame rate (for example, 60 frames per second) to generate an image sequence. Then input each image into the model constructed below.

[0069] Step S120, extracting text features F of the input image t :

[0070] The image is input into the open source visual language model CogVLM, and the corresponding text is generated using the prompt: "Describe the content of this image." Then the text information is input into the text encoder of CLIP to obtain the corresponding text feature F for each image. t . The process can be written as:

[0071] F t =Encoder(VLM(X in ,prompt)),

[0072] Among them, F t is the text feature, Encoder(·) is the text encoder of CLIP, VLM(·) is the open source visual language model CogVLM, X in is the input picture, and prompt is the prompt for the visual language model.

[0073] Step S130: construct a color retrieval enhancement gallery, input the original image for retrieval and obtain the color features F that need to be supplemented. c .

[0074] First, a Gaussian blur kernel is applied to the input image, which blurs the texture details but preserves the overall color distribution. Then, the embedding model ResNet is used to extract features from the image. This process can be written as:

[0075] G(X in )=GaussianBlur(X in ,kernel size ,σ),

[0076] F image =E(G(X in )),

[0077] Among them, X in is the input image, F image is the output image feature, E(·) is the model for extracting image features (embedding model), and G(·) is the Gaussian blur function applied to the input image with a kernel size of 33×33 and a standard deviation of 5.

[0078] Secondly, a vector database is established to retrieve the required color vector information corresponding to the input image. Specifically, the images in the ImageNet dataset are converted into feature vectors using the above embedding model and stored in the vector database.

[0079] Then, a search strategy is established to query the vector database.

[0080] Given a query F image , construct a small classifier C (a multilayer perceptron) to query the index of the most needed color vector. Set a hyperparameter to select the most needed k vectors. The whole process can be written as:

[0081] F c =CRAG(F image ,k,C),

[0082] Among them, F c is the output of the Color Retrieval Enhanced Gallery, and CRAG is the Color Retrieval Enhanced Gallery.

[0083] Step S140: the original input image X in , supplementary color feature information F c , the extracted text information F t Input DA 2 Net network: the supplementary color feature information F c , the extracted text information F t Add and original input image X in Input to DA 2 Net model, the process can be written as:

[0084] F in =F c +F t ,

[0085] F f =DA 2 Net(X in ,F in )

[0086] Among them, F f It's DA 2 Net network output.

[0087] in DA 2 The differential proxy attention mechanism is used in the local Transformer attention layer of the Net network to calculate the attention features: The differential proxy attention mechanism introduces a set of additional proxy tokens A into the traditional attention module, expressed as a four-tuple (Q, A, K, V). The proxy token A first acts as a proxy for the query token Q, responsible for aggregating information from K and V, and then broadcasting it back to Q, which can be expressed as:

[0088] O A =Attn S (Q,A,Attn S (A,K,V))=σ(QA T )σ(AK T )V

[0089] in, are the query, key, and value matrices, is the proxy token, and σ(·) represents the Softmax function. As shown in the above formula, the proxy attention mechanism consists of two attention operations. First, the proxy token A is regarded as a query, and attention calculation is performed between A, K, and V to aggregate the proxy feature V from all values. A Then, in the second attention calculation, A is used as the key and V A As the value, it works together with the query matrix Q to broadcast the global information of the proxy feature to each query tag, and finally obtains the output O A .

[0090] To calculate the first attention score, a gating mechanism is combined with a differential operation, which can be expressed as:

[0091] λ1=Sigmoid(β1)

[0092] σ(AK T )=λ1·Softmax(AK T )-(1-λ1)·Softmax(Conv(AK T ))

[0093] Among them, β1 is a learnable gating parameter, Sigmoid(·) is the activation function, and Conv(·) is the convolution operation. By applying a differential operation to two independently calculated Softmax attention maps and using a gating mechanism, this method reduces attention noise and achieves adaptive attention, allowing the model to focus more on key information.

[0094] In the second attention calculation, the global information from the agent features is assigned to each query token, expressed as:

[0095] λ2=Sigmoid(β2)

[0096] σ(QA T )=λ2·Softmax(QA T )-(1-λ2)·Softmax(Conv(QA T ))

[0097] Among them, β1 is a learnable gating parameter. Combined with the global information transfer in the second attention calculation, the model achieves a more robust and accurate attention distribution.

[0098] in DA 2 The local Transformer attention layer is used to extract features in the residual Transformer attention block of the Net network. The local Transformer attention layer consists of differential proxy attention, layer normalization and multi-layer perceptron. Given the input local feature X, the entire process in the local Transformer attention layer can be written as:

[0099] X=DAA(LN(X))+X,

[0100] X=MLP(LN(X))+X,

[0101] Where DAA(·) is the differential agent attention mechanism, LN(·) is layer normalization, and MLP(·) is a multi-layer perceptron consisting of two fully connected layers with a GELU activation function in between.

[0102] in DA 2 Extract and fuse features from the residual Transformer attention block of the Net network: Given the input feature F of the i-th residual Transformer attention block i,0 , first use L local Transformer attention layers (LTAL) to extract the intermediate features F i,1 ,F i,2 ,…,F i,L , the process can be expressed as:

[0103] F i,j =LTALi,j (F i,j-1 ),j=1,2,…,L

[0104] Among them, LTAL i,j (·) represents the jth local Transformer attention layer in the i-th residual Transformer attention block. Then, a convolutional layer is added before the residual connection. The output of the residual Transformer attention block can be expressed as:

[0105] F i,out =Conv i (F i,L )+F i,0

[0106] Among them, Conv i (·) is the convolutional layer in the i-th residual Transformer attention block.

[0107] Step S150, DA 2 Net network output and input image X in The final image restoration result is obtained by fusion

[0108] The restored image is obtained using the following formula:

[0109]

[0110] in, is the final image restoration result, F f It's DA 2 Net network output, X in is the input image, Conv(·) is the convolution operation, and the enhanced image sequence is reconstructed into a video. The enhanced image sequence is arranged in chronological order, the original resolution of the image is maintained, and the enhanced video is reconstructed at 60 frames per second to obtain a high-quality video output.

[0111] Traditional enhancement methods fail to effectively cope with dynamic scene changes in high-resolution videos, which can easily lead to inconsistent timing, flickering or color drift. Secondly, 4K videos usually contain a lot of detailed information, which places extremely high demands on computing power. Traditional methods are prone to computing bottlenecks when processing large-scale data. Finally, existing enhancement methods often ignore the semantic information in the video content, especially in complex lighting changes or scene transitions, which can easily lead to unnatural or incoherent enhancement effects.

[0112] To address these challenges, this paper proposes a novel approach. First, a Color Retrieval Augmentation Gallery (CRAG) is designed to enhance the color information of video frames by extracting color priors from a large number of natural images. Specifically, a large-scale Gaussian blur kernel is first applied to images in the ImageNet dataset to remove detailed textures in the images, and these processed images are stored in a gallery for retrieval.

[0113] Next, in the embedded color semantic space, the k word vectors that are most needed to supplement the color information of the input image are searched. This method effectively ensures the consistency of the color information of the generated image during the inference process, thereby obtaining a more natural and harmonious enhanced image. This process performs color correction and enhancement on each frame in the video to ensure color stability and visual consistency in the video stream.

[0114] Secondly, we introduce text prior by using a visual language model to extract text information from the input image. This text prior not only helps to avoid hallucination effects in the restoration of low dynamic range (LDR) images, but also enables accurate supplementation of details in the video based on contextual information. In 4K video enhancement, especially in fast scene changes or complex lighting conditions, text prior provides effective semantic guidance for dynamic scenes, reducing distortion or unnatural phenomena that may occur during the enhancement process.

[0115] Finally, the present invention proposes a differential proxy attention mechanism, which integrates a set of additional proxy tokens A into the traditional attention module, represented as a four-tuple (Q, A, K, V). The proxy token A acts as a proxy for the query token Q, first aggregating information from K and V, and then passing it back to Q. This method enhances the effectiveness of information transfer while maintaining the global context modeling capability of the model. In particular, when processing 4K videos, the temporal consistency between video frames is crucial. Through the differential operation, a gating mechanism is introduced when calculating the attention score, and the two independently calculated Softmax attention maps are differentially operated, thereby effectively reducing the noise in the timing and ensuring consistency and smoothness between frames. This method effectively alleviates the problem of timing inconsistency that may occur in the video enhancement process and avoids the typical flickering effect and loss of details.

[0116] Since video content usually has high dynamic variability and complex spatiotemporal information, this method can ensure that each frame can be accurately enhanced while ensuring the overall visual smoothness and color consistency of the video when processing 4K video through effective fusion of frame information and enhancement of multimodal information.

[0117] From the perspective of information extraction and fusion, the text features are obtained by using the visual language model to describe the image sequence, which utilizes the semantic information of the image. Traditional image and video enhancement methods often only focus on the visual features of the image itself, while this method introduces text features to understand the image content from a semantic level, providing richer information for enhancement, making the enhanced images and videos more in line with human understanding and expectations of the content.

[0118] Constructing a color retrieval enhancement gallery to obtain the color features that need to be supplemented is an effective way to mine and utilize the color information of images. Color plays a key role in the visual effects of images and videos. By supplementing color features, the color richness, vividness, and realism of images and videos can be significantly improved, making the enhanced images more visually attractive.

[0119] The image sequence, text features and color features are input into DA 2 Net model realizes the fusion of multiple features. This fusion method can give full play to the advantages of different features and comprehensively consider multiple factors such as image content, semantics and color, so that D A 2 Net model outputs more accurate and higher-quality results, laying the foundation for the subsequent generation of high-quality restored images and enhanced videos.

[0120] From the perspective of processing flow, extracting frames from the source video to obtain an image sequence is an efficient processing method. After decomposing the video into an image sequence, each frame can be processed in detail, avoiding the high computational complexity and high complexity of complex processing of the entire video, and also facilitating subsequent feature extraction and enhancement operations on the image.

[0121] Model combination and convolution fusion using DA 2 Net model, and then perform convolution operation on its output and fuse it with the original image sequence. This method combines the intelligent processing ability of the model with the local feature extraction ability of the convolution operation, and can effectively enhance the details and quality of the image while retaining the key information of the original image, and finally generate high-quality restored images and enhanced videos.

[0122] This method can achieve real-time enhancement of ultra-high-definition images and videos, which is of great significance in many practical application scenarios. For example, in the fields of live broadcasting and video surveillance, real-time enhancement can provide a clearer and higher-quality visual experience and meet users' demand for real-time picture quality.

[0123] This method has broad application prospects. The real-time enhancement technology of ultra-high-definition images and videos can be applied to multiple fields, such as film and television production, virtual reality, medical imaging, etc. In film and television production, it can improve the visual effects of the film; in virtual reality, it can enhance the sense of immersion; in medical imaging, it helps doctors observe lesions and other details more clearly, and has broad application prospects and commercial value.

[0124] Reference Figure 3 , Figure 3 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by an embodiment of the present invention; the method for real-time enhancement of ultra-high-definition images and videos includes but is not limited to the following steps:

[0125] Step S210, inputting the image sequence into a visual language model;

[0126] Step S220, inputting a prompt, and the visual language model performs a text description of the image sequence based on the prompt;

[0127] Step S230: compile the text description of the image sequence by a text compiler to obtain text features.

[0128] The image sequence is input into the visual language model, and the image sequence is described in text by the visual language model to obtain the corresponding text features of each image, including the following steps: inputting the image sequence into the visual language model; inputting the prompt, and the visual language model performs a text description of the image sequence based on the prompt; and compiling the text description of the image sequence through the text compiler to obtain the text features. The image is input into the open source visual language model CogVLM, and the corresponding text is generated using the prompt: "Describe the content of this image." Then the text information is input into the text encoder of CLIP to obtain the corresponding text features F of each image. t . The process can be written as:

[0129] F t =Encoder(VLM(X in ,prompt)),

[0130] Among them, F t is the text feature, Encoder(·) is the text encoder of CLIP, VLM(·) is the open source visual language model CogVLM, X in is the input picture, and prompt is the prompt for the visual language model.

[0131] Inputting image sequences and prompts into the visual language model provides a clear task orientation for the model's processing. Prompts can guide the visual language model to focus on specific information in the image sequence, such as the image's theme, scene, and object attributes, making the text description generated by the model more targeted and relevant, and better reflecting the key information of the image, thereby providing more valuable semantic guidance for subsequent image enhancement.

[0132] By changing the prompts, the visual language model can flexibly adjust the way it understands and describes the image sequence. Different application scenarios and needs may require attention to different aspects of the image. For example, in artistic creation, the image style and emotional expression may be more concerned, while in security monitoring, the target object and behavior in the image may be more concerned. Using prompts can easily adapt to these different needs, making the model more versatile and adaptable.

[0133] The visual language model has a strong cross-modal understanding capability and can perform in-depth semantic analysis and understanding of image sequences. It can convert the visual information in the image into natural language descriptions and mine the implicit semantic relationships and contextual information in the image. Compared with traditional methods based only on image visual features, this text description based on semantic understanding can provide richer and more abstract information, help better grasp the content and meaning of the image, and provide a more comprehensive basis for image enhancement.

[0134] The text description generated by the visual language model can present various information about the image in the form of natural language. Natural language has rich vocabulary and grammatical structure, which can describe the details and features in the image more accurately and meticulously. This rich information can provide more dimensional references for subsequent image enhancement. For example, in tasks such as image restoration and super-resolution reconstruction, specific parts of the image can be restored or enhanced based on the information in the text description, making the enhanced image more in line with human cognition and expectations.

[0135] The text features are obtained by compiling the text description of the image sequence through a text compiler. This process refines and standardizes the text information. The text compiler can remove redundant information in the text description, extract key semantic features, and convert them into a format suitable for subsequent processing. The text features obtained in this way are more concise and accurate, which is conducive to improving the processing efficiency and accuracy of the model. At the same time, it is also easy to merge with other types of features (such as image features and color features), providing more effective support for real-time enhancement of images and videos.

[0136] The compiled text features are presented in a structured form, which is easier to be used by subsequent models (such as D 2Net model). This structured representation enables the model to better understand and utilize text information, thereby more effectively combining semantic information in the image and video enhancement process, improving the quality and accuracy of the enhancement effect.

[0137] Reference Figure 4 , Figure 4 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by an embodiment of the present invention; the method for real-time enhancement of ultra-high-definition images and videos includes but is not limited to the following steps:

[0138] Step S310, applying a Gaussian blur kernel to the image sequence to blur texture details and preserve the overall color distribution;

[0139] Step S320, extracting image feature vectors from the image sequence using the embedding model ResNet;

[0140] Step S330: storing the image feature vector into a vector database to obtain a color retrieval enhancement gallery.

[0141] A Gaussian blur kernel is applied to the input image, which blurs the texture details but preserves the overall color distribution. Then, the embedding model ResNet is used to extract features from the image. This process can be written as:

[0142] G(X in )=GaussianBlur(X in ,kernel size ,σ),

[0143] F image =E(G(X in )),

[0144] Among them, X in is the input image, F image is the output image feature, E(·) is the model for extracting image features (embedding model), and G(·) is the Gaussian blur function applied to the input image with a kernel size of 33×33 and a standard deviation of 5.

[0145] A vector database is built to retrieve the required color vector information corresponding to the input image. The images in the ImageNet dataset are converted into feature vectors using the above embedding model and stored in the vector database, thus obtaining the color retrieval enhancement gallery.

[0146] Applying a Gaussian blur kernel to an image sequence blurs texture details and retains the overall color distribution. This operation is highly targeted. In many image enhancement tasks, texture details may interfere with the extraction and analysis of color features. Through Gaussian blur processing, these unnecessary detail information is removed or weakened, so that subsequent processing can focus more on the color features of the image, improving the accuracy and efficiency of color feature extraction. After Gaussian blurring, the complexity of the image is reduced, reducing the redundant information in the data volume. This not only helps to speed up the subsequent feature extraction and storage, but also reduces the consumption of computing resources, making the entire system more efficient and stable when processing large-scale image sequences.

[0147] The present invention uses the embedding model ResNet to extract image feature vectors from image sequences. ResNet is a deep residual network with excellent feature extraction capabilities, which can effectively extract rich and representative features from images. In the process of building a color retrieval enhanced gallery, it can accurately capture the color features of the image and other potential features related to color, providing more comprehensive and accurate information for subsequent color retrieval.

[0148] The present invention can store the extracted image feature vectors into a vector database to obtain a color retrieval enhancement gallery, which facilitates rapid retrieval and matching. The vector database has an efficient storage and retrieval algorithm, which can quickly find other vectors similar to the input image feature vector, thereby realizing rapid retrieval of color features. This is crucial for ultra-high-definition image and video real-time enhancement tasks with high real-time requirements, and can obtain the color features that need to be supplemented in a short time, thereby improving the efficiency of the entire enhancement process.

[0149] The present invention uses a vector database to store image feature vectors, so that the color retrieval enhancement gallery has good scalability and flexibility. New image feature vectors can be easily added to continuously enrich the content of the gallery to adapt to different image and video data. At the same time, by adjusting the retrieval algorithm and parameters of the vector database, the retrieval accuracy and range of color features can be flexibly controlled according to specific application requirements.

[0150] Reference Figure 5 , Figure 5 A flowchart of a method for real-time enhancement of ultra-high-definition images and videos provided by an embodiment of the present invention; the method for real-time enhancement of ultra-high-definition images and videos includes but is not limited to the following steps:

[0151] Step S410, given an image feature vector for query, construct a classifier to query the index of the most needed color vector;

[0152] Step S420 , selecting the most needed k word vectors, and searching in the color retrieval enhancement gallery based on the image feature vector and the classifier to obtain the color features that need to be supplemented.

[0153] The image sequence is input into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented, including the steps of: given an image feature vector for query, constructing a classifier to query the index of the most needed color vector; selecting the most needed k word vectors, and retrieving the color features that need to be supplemented in the color retrieval enhancement gallery based on the image feature vector and the classifier.

[0154] Establish a retrieval strategy to query the vector database: Given a query F image , we construct a small classifier C (a multilayer perceptron) to query the index of the most needed color vector. We set a hyperparameter to select the most needed k vectors, and the whole process can be written as:

[0155] F c =CRAG(F image ,k,C),

[0156] Among them, F c is the output of the Color Retrieval Enhanced Gallery, and CRAG is the Color Retrieval Enhanced Gallery.

[0157] The present invention can give an image feature vector for query and construct a classifier to query the most needed color vector index. This method focuses the retrieval target on the color vector closely related to the current image features, and can accurately locate the part of the image that may be missing or needs to be enhanced in terms of color. Compared with the aimless retrieval method, it greatly improves the retrieval efficiency and accuracy, makes the obtained color features more targeted, and can directly meet the image enhancement's demand for color information.

[0158] The invention provides flexibility in selecting the most needed k word vectors. By adjusting the k value, the number of retrieved color features can be flexibly controlled according to the specific image enhancement requirements and computing resource limitations. If richer and more detailed color information is needed to achieve a more refined enhancement effect, the k value can be appropriately increased; if the computing resources are more sensitive, or only the image color needs to be roughly adjusted, the k value can be reduced. This adjustability enables the method to adapt to different application scenarios and requirements.

[0159] The present invention can search in the color retrieval enhancement gallery based on the image feature vector and the classifier, and can make full use of the large amount of image feature vector information stored in the gallery. The color retrieval enhancement gallery itself is carefully constructed and contains rich color feature data. Through this retrieval method, the color feature that best matches the current image can be mined from the gallery, providing strong support for image enhancement. This process is equivalent to accurate screening in a huge color feature library, which greatly increases the possibility of obtaining suitable color features.

[0160] The present invention can use image feature vectors and classifiers for retrieval, ensuring that the acquired color features have a high correlation with the original image features. This correlation enables the supplemented color features to be better integrated into the original image, while enhancing the image color effect and maintaining the overall consistency and coordination of the image. For example, when color enhancement is performed on a scene in an ultra-high-definition video, the retrieved color features can match the image content of the scene, and there will be no abrupt or inharmonious colors, thereby improving the visual quality of the entire video.

[0161] The present invention can accurately obtain the color features that need to be supplemented, providing key information support for subsequent image and video enhancement. By combining these color features with other features (such as image features and text features), images and videos can be processed more comprehensively, thereby significantly improving the real-time enhancement effect of ultra-high-definition images and videos. Whether in terms of color richness, vividness or fit with image content, better performance can be achieved, bringing users a better visual experience.

[0162] This retrieval method has good scalability and versatility. With the continuous enrichment and updating of image feature vectors in the color retrieval enhancement gallery, the retrieved color features will be more comprehensive and accurate. At the same time, this retrieval method based on feature vectors and classifiers is not limited to specific image or video datasets, but can be applied to various types of image and video enhancement tasks, and has broad application prospects.

[0163] In some embodiments, the method for real-time enhancement of ultra-high-definition images and videos includes the following steps: performing a video frame extraction operation on a source video to obtain a corresponding image sequence; inputting the image sequence into a visual language model, and performing an image text description on the image sequence through the visual language model to obtain the corresponding text features of each image; constructing a color retrieval enhancement gallery, and inputting the image sequence into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented; inputting the image sequence, text features, and color features into the D 2 Net model, and obtain DA 2 Net model output; 2The output of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0164] DA 2 The Net model includes: multiple residual Transformer attention blocks and convolutional layers. The residual Transformer attention blocks are used to extract and fuse features. The residual Transformer attention blocks include multiple local Transformer attention layers and convolutional layers. The local Transformer attention layers include: differential proxy attention mechanism, layer normalization and multi-layer perceptron. The differential proxy attention mechanism is used to calculate attention features.

[0165] Differential proxy attention mechanism: A set of additional proxy tokens A is introduced into the traditional attention module, represented as a four-tuple (Q, A, K, V). The proxy token A first acts as a proxy for the query token Q, responsible for aggregating information from K and V, and then broadcasting it back to Q, expressed as:

[0166] O A =Attn S (Q,A,Attn S (A,K,V))=σ(QA T )σ(AK T )V

[0167] in, are the query, key, and value matrices, is the proxy token and σ(·) represents the Softmax function.

[0168] Extracting and fusing features in the residual Transformer attention block includes the following steps:

[0169] Given the input feature F of the i-th residual Transformer attention block i,0 , first use L local Transformer attention layers (LTAL) to extract the intermediate features F i,1 ,F i,2 ,…,F i,L , the process is expressed as:

[0170] F i,j =LTAL i,j (F i,j-1 ),j=1,2,…,L

[0171] Among them, LTAL i,j(·) represents the jth local Transformer attention layer in the i-th residual Transformer attention block. Then, a convolutional layer is added before the residual connection. The output of the residual Transformer attention block is expressed as:

[0172] F i,out =Conv i (F i,L )+F i,0

[0173] Among them, Conv i (·) is the convolutional layer in the i-th residual Transformer attention block.

[0174] to DA 2 The output of the Net model is convolved and fused with the image sequence to obtain the restored image, including the following steps:

[0175] The restored image is obtained using the following formula:

[0176]

[0177] in, is the final restored image, F f It's DA 2 The output of the Net model, X in is an image sequence, and Conv(·) is a convolution operation.

[0178] Improved results:

[0179] First, we verify the effectiveness of our model in the image field. Compared with existing image enhancement methods (such as HDRNet, ZeroDCE, RUAS, LCDPNet, RetinexFormer, Sagiri, CSEC, CECF), our method is experimentally verified in the LCDP dataset and Mobile-Spec dataset, which include overexposed and underexposed images, as well as the underwater image SUID dataset.

[0180] The experiments were conducted on a single NVIDIA A100 80GB GPU. 2 Net, the number of residual Transformer attention blocks is set to 6, and each residual Transformer attention block contains 6 local Transformer attention layers. The embedding dimension is set to 66, and the number of attention heads is set to 6.

[0181] The performance is evaluated using three widely recognized metrics: PSNR, SSIM, and LPIPS. PSNR and SSIM focus on evaluating the fidelity and structural integrity of the reconstructed image, while LPIPS evaluates the perceptual quality by measuring the similarity of deep features, providing an evaluation that is more in line with human visual quality.

[0182] Among them, PSNR (Peak Signal-to-Noise Ratio) is a common indicator for measuring image quality. It is mainly used to evaluate the difference between the processed image and the original image. The larger the value, the better the image quality. The calculation formula is:

[0183]

[0184] Among them, MSE stands for mean square error, which is used to measure the difference between two images at the pixel level. I(i, j) and K(i, j) respectively represent the pixel values ​​of the generated fused image and the corresponding gold standard image (ground truth) at position (i, j). m and n respectively represent the width and height of the image. MAX represents the maximum value of the image pixels.

[0185] SSIM (Structural Similarity Index Measure) is an indicator used to evaluate the quality of two images. It mainly measures the similarity of images by comparing the brightness, contrast and structural information of the images. The closer the value is to 1, the better the image quality.

[0186] The calculation formula of SSIM is:

[0187] c1=(k1L) 2 ;

[0188] c2=(k2L) 2 ;

[0189]

[0190] Among them, μ Y is the average value of the fused image Y output by the model, is the gold standard Y corresponding to Y gt The average value of is the image Y and Y gt The covariance of is the variance of Y, Yes gt variance; L is the dynamic range of pixel values, k1 and k2 represent preset hyperparameters, where k1=0.01, k2=0.03, c1 and c2 represent smoothing parameters.

[0191] The LPIPS (Learned Perceptual Image Patch Similarity) metric is a metric used to evaluate the perceptual similarity between two images. It is based on features extracted by a deep convolutional neural network rather than simple pixel differences. The smaller the value, the better the image quality.

[0192] The calculation formula of LPIPS is:

[0193]

[0194] Among them, x represents the fused image output by the model, x0 represents the gold standard image corresponding to x, and f l (x) and f l (x0) represents the lth layer feature map extracted by a pre-trained VGG deep neural network. l (x) is the feature of image x at layer l, f l (x0) is the feature of image x0 at layer l. represents the Euclidean distance (L2 norm) between feature maps. This is the feature map f l (x) and f l The difference metric between (x0) is used to measure the similarity of the two images in the lth layer features. l Represents the weight coefficient of each layer feature map. Different network layers have different effects on perceived similarity, and the weight w l Used to reflect this, it is usually learned through training. It means summing the results of all layers l and comprehensively considering the feature differences at different levels.

[0195] Table 1 Results on LCDP dataset

[0196]

[0197] Table 2 Results on the Mobile-Spec dataset

[0198]

[0199] Table 3 Results on the SUID dataset

[0200]

[0201]

[0202] Table 1 lists the comparison results on the LCDP dataset, and all methods are retrained under the same settings. 2Net outperforms the second best method, RetinexFormer, by about 0.36dB and 0.02dB higher in PSNR and SSIM, respectively, while reducing LPIPS by 0.03.

[0203] Experiments were conducted on the Mobile-Spec dataset. As shown in Table 2, this method achieved the best performance on all evaluation indicators. In addition, Table 3 shows that DA 2 Net can effectively correct the color of underwater images, while existing image enhancement methods perform poorly in underwater scenes. Specifically, compared with the second best method, our method improves by 4.52dB in PSNR, 0.05dB in SSIM, and reduces by 0.0027 in LPIPS.

[0204] Overall, this method shows excellent results in enhancing overexposed, underexposed and underwater images, and significantly improves the image quality in various complex scenes.

[0205] In addition, an experimental evaluation was conducted on 4K low-quality videos. The experimental results show that in the inference stage, this method can process ultra-high-resolution 4K videos at an efficient speed. Specifically, the frame rate of this method when processing 4K videos reaches 68 frames per second (FPS), which can significantly improve the visual effect of the video and enhance the image details and color performance even in the case of low-quality videos.

[0206] In one embodiment, the ultra-high-definition image and video real-time enhancement system includes: a video frame extraction module, which is used to perform a video frame extraction operation on a source video to obtain a corresponding image sequence; a text feature extraction module, which is connected to the video frame extraction module in communication, and is used to input the image sequence into a visual language model, and perform an image text description on the image sequence through the visual language model to obtain the corresponding text features of each image; a color feature extraction module, which is connected to the video frame extraction module in communication, constructs a color retrieval enhancement gallery, and inputs the image sequence into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented; 2 Net model, inputting image sequences, text features, and color features into the D 2 Net model, and obtain DA 2 Net model output; video enhancement module, and DA 2 Net model communication connection, to DA 2 The output of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

[0207] The processor and the memory may be connected via a bus or other means.

[0208] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0209] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, and may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0210] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor or controller, for example, by a processor in the above-mentioned terminal embodiment, so that the above-mentioned processor can execute the ultra-high-definition image and video real-time enhancement method in the above-mentioned embodiment.

[0211] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0212] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present invention. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.

[0213] The specific implementation of the present invention described above does not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and modifications made based on the technical concept of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A method for real-time enhancement of ultra-high-definition images and videos, characterized in that: The following steps are involved: Performing a video frame extraction operation on the source video to obtain a corresponding image sequence; Inputting the image sequence into a visual language model, and performing image text description on the image sequence through the visual language model to obtain text features corresponding to each image; Constructing a color retrieval enhancement gallery, inputting the image sequence into the color retrieval enhancement gallery to retrieve and obtain the color features that need to be supplemented; The image sequence, the text features, and the color features are input into the D 2 Net model, and obtain the DA 2 The output of Net model; The DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

2. The method for real-time enhancement of ultra-high-definition images and videos according to claim 1, characterized in that: The image sequence is input into a visual language model, and image text description is performed on the image sequence by the visual language model to obtain text features corresponding to each image, including the steps of: inputting the image sequence into a visual language model; Input a prompt, and the visual language model performs a text description on the image sequence based on the prompt; The text description of the image sequence is compiled by a text compiler to obtain text features.

3. The method for real-time enhancement of ultra-high-definition images and videos according to claim 1, characterized in that: Constructing a color retrieval enhancement gallery includes the following steps: Applying a Gaussian blur kernel to the image sequence to blur texture details and preserve overall color distribution; Extracting image feature vectors from the image sequence using an embedding model ResNet; The image feature vector is stored in a vector database to obtain the color retrieval enhanced gallery.

4. The method for real-time enhancement of ultra-high-definition images and videos according to claim 3, characterized in that: Inputting the image sequence into the color retrieval enhancement gallery to retrieve and obtain the color features that need to be supplemented, including the steps of: Given a feature vector of the image for query, construct a classifier to query the index of the most desired color vector; The most needed k word vectors are selected, and based on the image feature vector and the classifier, the color retrieval enhancement gallery is searched to obtain the color features that need to be supplemented.

5. The method for real-time enhancement of ultra-high-definition images and videos according to claim 1, characterized in that: The DA 2 Net models include: Multiple residual Transformer attention blocks and convolutional layers. The residual Transformer attention blocks are used to extract and fuse features. The residual Transformer attention block includes multiple local Transformer attention layers and convolutional layers; The local Transformer attention layer includes: a differential proxy attention mechanism, layer normalization and a multi-layer perceptron, and the differential proxy attention mechanism is used to calculate attention features.

6. The method for real-time enhancement of ultra-high-definition images and videos according to claim 5, characterized in that: The differential proxy attention mechanism introduces a set of additional proxy tokens A into the traditional attention module, represented as a four-tuple (Q, A, K, V). The proxy token A first acts as a proxy for the query token Q, responsible for aggregating information from K and V, and then broadcasting it back to Q, expressed as: Among them, Q, K, are the query, key, and value matrices, is the proxy token and σ(·) represents the Softmax function.

7. The method for real-time enhancement of ultra-high-definition images and videos according to claim 5, characterized in that: Extracting and fusing features in the residual Transformer attention block includes the following steps: Given the input feature F of the i-th residual Transformer attention block i,0 , first use L local Transformer attention layers (LTAL) to extract the intermediate features F i,1 ,F i,2 ,…,F i,L , the process is expressed as: F i,j =LTAL i,j (F i,j-1 ),j=1,2,…,L Among them, LTAL i,j (·) represents the jth local Transformer attention layer in the i-th residual Transformer attention block. Then, a convolutional layer is added before the residual connection. The output of the residual Transformer attention block is expressed as: F i,out =Conv i (F i,L )+F i,0 Among them, Conv i (·) is the convolutional layer in the i-th residual Transformer attention block.

8. The method for real-time enhancement of ultra-high-definition images and videos according to claim 1, characterized in that: The DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, including the steps of: The restored image is obtained using the following formula: in, is the final restored image, F f It's DA 2 The output of the Net model, X in is the image sequence, and Conv(·) is the convolution operation.

9. A real-time ultra-high-definition image and video enhancement system, characterized in that: include: A video frame extraction module is used to perform a video frame extraction operation on the source video to obtain a corresponding image sequence; A text feature extraction module, which is in communication with the video frame extraction module and is used to input the image sequence into a visual language model, and to perform image text description on the image sequence through the visual language model to obtain text features corresponding to each image; A color feature extraction module is connected to the video frame extraction module for constructing a color retrieval enhancement gallery, and inputs the image sequence into the color retrieval enhancement gallery for retrieval to obtain the color features that need to be supplemented; DA 2 Net model, inputting the image sequence, the text features, and the color features into the D 2 Net model, and obtain the DA 2 The output of Net model; Video enhancement module, with the DA 2 Net model communication connection, the DA 2 The output result of the Net model is convolved and fused with the image sequence to obtain a restored image, and an enhanced video is generated based on the restored image.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the ultra-high-definition image and video real-time enhancement method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Short video popularity prediction method based on multi-modal retrieval enhancement

    CN119172573A

  • System and method for adaptive video fast forward using scene generative models

    US20040175058A1

Cited By

  • Underwater image enhancement method and system based on vision-text fusion

    CN120634934A

  • A brand promotion video gray frame enhancement method and system

    CN122597183A