Transform-based high-power face video super-resolution processing method

By constructing a Transformer-based high-magnification face video super-resolution technology framework, and utilizing multi-channel information extraction, inter-frame point-by-point mask attention, and low-rank decomposition to optimize the attention mechanism, the problems of information missing and feature drift in high-magnification face video super-resolution are solved, achieving high-quality and efficient face video reconstruction.

CN120765460AActive Publication Date: 2025-10-10BEIJING UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510834600.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-10
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing technologies suffer from information loss and feature drift in high-magnification face video super-resolution tasks, resulting in insufficient and unstable reconstruction quality.

Method used

A Transformer-based high-magnification face video super-resolution technology framework is constructed, including a multi-channel information extraction module, an inter-frame point-by-point mask attention module, and a low-rank decomposition attention mechanism parameter optimization method. Through multi-dimensional feature representation, temporal consistency, and computational parameter optimization, the inter-frame feature consistency and computational efficiency are improved.

Benefits of technology

Effectively restore the details of high-magnification facial videos, improve reconstruction quality and stability, reduce computational complexity, and achieve higher-quality facial image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765460A_ABST
    Figure CN120765460A_ABST
Patent Text Reader

Abstract

The invention discloses a high-power face video super-resolution processing method based on Transform, and belongs to the technical field of video super-resolution. Comprising the following contents: firstly, a multi-channel-oriented multi-dimensional feature representation method; and 2, a time uniformity method based on an inter-frame point-by-point mask attention module. And thirdly, an attention mechanism parameter optimization method based on low-rank decomposition. According to the method, the pertinence of information processing in the video frames and among the video frames is fully utilized and combined with the thought of time-space consistency of the video super-division task, efficient reconstruction of the high-power face video frames is achieved, and the overall super-division quality of the video is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video super-resolution, and in particular relates to a high-magnification face video super-resolution processing method based on Transformer. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, research in the field of video processing has made significant progress. In particular, video super-resolution (VSR) technology, which improves video quality and detail, has become a hot topic. Given the increasing use of facial images in scenarios such as social media, surveillance, and video communications, facial video super-resolution technology is particularly important.

[0003] Traditional video super-resolution methods rely heavily on interpolation techniques and filtering algorithms. While these methods can improve video quality to a certain extent, they often struggle to recover high-frequency details and are limited in their effectiveness when processing dynamic scenes. In recent years, super-resolution techniques based on deep learning have gradually emerged. Using models such as convolutional neural networks (CNNs), they can effectively capture complex features and structural information in images, enabling more refined image enhancement.

[0004] However, existing deep learning methods still face technical challenges when it comes to super-resolution of high-resolution facial videos. First, traditional CNNs are prone to information loss and blurring when processing long video sequences, especially when recovering details. Furthermore, the temporal nature of video data requires super-resolution techniques to not only process single frames but also consider the temporal correlation between frames to maintain video clarity and feature consistency.

[0005] To address these issues, Transformer-based architectures have been introduced to video super-resolution research in recent years. The Transformer's self-attention mechanism effectively captures long-range dependencies, making it suitable for processing time-series data. Applying the Transformer to face video super-resolution tasks can enhance spatial detail while capturing temporal information between frames. This approach not only improves the super-resolution effect of a single frame but also achieves higher clarity and feature consistency across video sequences, enhancing overall video quality.

[0006] For the standard Transformer architecture, computational complexity has always been a pain point in the process of processing long sequence problems. Through analysis of computational complexity, this paper proposes an attention mechanism parameter optimization method based on low-rank decomposition, which can effectively alleviate the computational complexity problem of the model.

[0007] Based on the attention mechanism based on low-rank decomposition technology, the present invention constructs a framework for high-magnification face video super-resolution technology.

[0008] First, a multi-channel information extraction module is constructed to obtain multi-channel spatial information. Secondly, an inter-frame point-by-point mask attention module is constructed to obtain the inter-frame motion trajectory. Then, a frame reconstruction module based on low-rank decomposition is used to integrate the time and space dimensions. Finally, the proposed Gan loss, inter-frame point-by-point loss and mean square error loss are used to improve the detail recovery and generation capabilities of high-magnification face videos, achieving more realistic and clear face image reconstruction.

[0009] This novel technical architecture can effectively promote the development of facial video super-resolution technology and provide higher-quality facial video processing solutions for various application scenarios. Summary of the Invention

[0010] The technical problem to be solved by this invention is the insufficient and unstable reconstruction quality caused by large-scale information loss and feature drift in high-magnification super-resolution (6, 8, 10) facial videos. To this end, a Transformer-based high-magnification facial video super-resolution technology framework is proposed. First, a multi-channel information extraction module is constructed to extract the spatial information contained in the input video frames to ensure the diversity of spatial information. Then, an inter-frame point-by-point mask attention module is used to integrate the motion trajectories between different frames, obtain temporal information, and strengthen the consistency of inter-frame features. Finally, a low-rank decomposition-based frame reconstruction module is used to integrate spatiotemporal information while reducing the computational parameters of the attention mechanism and improving the inference speed.

[0011] To achieve the above object, the present invention adopts the following technical solutions:

[0012] The Transformer-based high-resolution face video super-resolution technology framework includes the following:

[0013] First: multi-dimensional feature representation method for multiple channels;

[0014] In the face video super-resolution task, the present invention inputs five adjacent frames, denoted as n-2, n-1, n, n+1, and n+2. The size of each frame is typically represented using three dimensions: channel, height, and weight. For the 8x face video super-resolution task, the input frame size is 3×64×64, where 3 represents the channel, 64 represents the height, and 64 represents the width. The output is the predicted intermediate frame np, which is 3×512×512 in size.

[0015] The input frame is first expanded to 3×256×256 using bilinear interpolation and data augmentation techniques. Then, a vector transformation operation is performed to change the input frame representation to 48×64×64. The input frame is then split into four 12×64×64 frames, two 24×64×64 frames, and one 48×64×64 frame, based on the channel dimension. Convolution is used to expand each 12×64×64 frame to 64×64×64. An attention mechanism is used to build long-range dependencies between features in each 64×64×64 frame. Each 64×64×64 frame is then concatenated to form a 256×64×64 frame. Similar operations are performed on the two 24×64×64 frames and the one 48×64×64 frame, resulting in an output of 256×64×64, as shown in Equation 1.

[0016] Channel_Enhance(n+i)=λUpsample(n+i)+Attention (1)

[0017] Among them, i takes values ​​of -2, -1, 0, 1, 2, indicating the current input frame position, λ takes values ​​of 4, 2, 1, indicating the number of samples divided by channel, Upsample represents the upsampling method of the present invention, and Attention represents the attention mechanism after parameter optimization of the present invention.

[0018] Finally, the three groups of 256×64×64 are averaged to effectively fuse information from different sources and generate a representative output, which provides sufficient spatial information support for the subsequent inter-frame point-by-point mask attention module and frame reconstruction module.

[0019] Second: a temporal consistency method based on an inter-frame point-by-point mask attention module;

[0020] In the face video super-resolution task, five consecutive image frames are input, denoted as n-2, n-1, n, n+1, and n+2, and the output is the predicted intermediate frame, denoted as np. Processing according to Equation 1 yields enhanced information for the five frames. However, integrating this information into the predicted intermediate frame np is a key issue in face video super-resolution. To address this issue, a temporal consistency method based on an inter-frame point-by-point masked attention module was designed, as shown in Equations 2 and 3.

[0021] Difference(i)=Pixel(n+i)-Pixel(n) (2)

[0022] Mask:Pixel(n+i)-Pixel(n)≤0.00001 (3)

[0023] Among them, Formula 2 is the calculation method of pixel differences between frames, and Formula 3 is the Mask identification standard.

[0024] In the standard Transformer architecture, the interactive attention module takes as input a query, a key, and a value. This paper uses the information from the input intermediate frame n as the query and calculates the pixel differences between frames n-2, n-1, n+1, and n+2 and frame n. Positions where the absolute value of the pixel difference is less than 0.00001 are masked to guide the model's focus on point-by-point information between frames, strengthening attention distribution.

[0025] Third: Attention mechanism parameter optimization method based on low-rank decomposition;

[0026] In the self-attention mechanism, the input is X, and the query (Query), key (Key)

[0027] The sum value matrix calculates the output. The calculation formula of the attention mechanism is:

[0028]

[0029] In the standard attention mechanism calculation, Q is equal to the product of the input X and the parameter W, and W is subjected to a low-rank decomposition technique such that:

[0030]

[0031] in,

[0032] W∈R embed_dim×embed_dim ,(5)

[0033]

[0034] W low ∈R rank×embed_dim ,

[0035] In this paper, the size of embed_dim is 256, and the rank is 64. Therefore, in the standard attention mechanism, the parameter size corresponding to Q is the product of 256 and 256, which is 65536. After using the low-rank decomposition technique, the parameter size corresponding to Q is the product of 2, 64 and 256, which is 32768, which is half the number of parameters. Similarly, after low-rank decomposition is performed on the parameters corresponding to K and V, their parameter sizes are also halved.

[0036] By adjusting the rank, the parameter reduction multiple can be freely controlled. The specific reduction scale can be calculated using Formula 6.

[0037] Scale=embed_dim÷(rank×2) (6)

[0038] The attention mechanism parameter optimization method based on low-rank decomposition can provide different parameter compression ratios to meet the actual application computing needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is the overall framework of Transformer-based high-magnification face video super-resolution technology.

[0040] Figure 2 It is a temporal consistency method based on inter-frame point-by-point mask attention module.

[0041] Figure 3 A parameter optimization method for the attention mechanism based on low-rank decomposition. DETAILED DESCRIPTION

[0042] The present invention mainly implements a high-resolution face video technology framework based on Transformer. The specific method adopted by the present invention will be described in detail below with reference to the accompanying drawings.

[0043] Specifically, the overall framework of Transformer-based high-resolution face video super-resolution technology is as follows: Figure 1 As shown, it includes the following aspects: First: a multi-dimensional feature representation method for multiple channels. Second: a temporal consistency method based on an inter-frame point-by-point mask attention module. Third:

[0044] Attention mechanism parameter optimization method based on low-rank decomposition.

[0045] First: Multi-channel oriented multi-dimensional feature representation method:

[0046] Step 1: Input the video frame sequence and read the parameter file;

[0047] Step 2: Segment the information in the channel dimension according to Formula 1, and perform different upsampling processing on the information in different channels;

[0048] Step 3: Use the attention mechanism module optimized by Formula 5 to perform data enhancement on the upsampled information.

[0049] Step 4: Merge the data processed and enhanced according to multi-channel.

[0050] Second: Temporal consistency method based on inter-frame point-by-point mask attention module:

[0051] Step 5: Calculate the inter-frame point-by-point information using Formula 2;

[0052] Step 6: Mask the pixels that meet the conditions according to formula 3.

[0053] Third: Attention mechanism parameter optimization method based on low-rank decomposition:

[0054] Step 7: According to formula 5, the parameter calculation in formula 4 is improved, and according to the analysis in formula 6, the rank value is appropriately adjusted, the parameter scale is controlled, and in the experiment, the rank value is 64;

[0055] Step 8: A frame reconstruction module is constructed by a parameter optimization method based on a low-rank decomposition attention mechanism to obtain a high-quality intermediate frame np.

[0056] The above specific embodiments are only used to illustrate the technical solutions of the present application, and not to limit it. Those skilled in the art should understand that the above embodiments do not limit the present application in any form, and any similar technical solutions obtained by equivalent replacement or equivalent transformation, etc. belong to the protection scope of the present application.

Claims

1. A high-resolution face video processing method based on Transformer, characterized by: The following steps are involved: Step 1: Multi-channel multi-dimensional feature representation; In the face video super-resolution task, 5 adjacent frames are input, denoted as n-2, n-1, n, n+1, and n+2. The size of each frame is represented by three dimensions: channel, height, and width. Step 2: Temporal consistency based on the point-by-point mask attention module between frames; In the face video super-resolution task, 5 consecutive frames of images are input, denoted as n-2, n-1, n, n+1, n+2, and the output is the predicted intermediate frame, denoted as np; Step 3: Optimize the parameters of the attention mechanism based on low-rank decomposition; In the self-attention mechanism, the input is X, and the output is calculated by querying Q, key K, and value V matrices. In the interactive attention module in the standard Transformer architecture, the input includes query, key, and value. The information of the input intermediate frame n is used as the query, and the pixel difference between the n-2 frame, n-1 frame, n+1 frame, and n+2 frame is calculated respectively. For positions where the absolute value of the pixel difference is less than 0.00001, a mask is marked to guide the model to focus on the point-by-point information between frames and strengthen the attention distribution.

2. The Transformer-based high-resolution face video processing method according to claim 1, characterized in that: In step 1, for the 8x face video super-resolution task, the size of the input frame is 3×64×64, where 3 represents the channel, 64 represents the height and 64 represents the width, and the output is the predicted intermediate frame np, with a size of 3×512×512; For the input frame, it is first expanded to 3×256×256 through bilinear interpolation and data enhancement, and then the representation of the input frame is changed to 48×64×64 through vector transformation operation; according to the channel dimension, the input frame is divided into 4 12×64×64, 2 24×64×64, and 1 48×64×64; convolution is used to expand the channel of each 12×64×64 to 64×64×64, and the attention mechanism is used to build the long-range dependency of the features in each 64×64×64, and then each 64×64×64 is spliced ​​into 256×64×64; similar operations are performed on the two 24×64×64 and one 48×64×64, and the output of 256×64×64 is also obtained, as shown in formula (1); Channel_Enhance(n+i)=λUpsample(n+i)+Attention (1) Among them, Channel_Enhance represents the enhanced channel information, n represents the input intermediate frame, i takes values ​​of -2, -1, 0, 1, and 2, representing the current input frame position, representing the first two frames, the first frame, the current frame, the next frame, and the next two frames of the intermediate frame respectively; λ takes values ​​of 4, 2, and 1, representing the number of samples divided by channel; Upsample represents the upsampling method; Attention represents the attention mechanism after parameter optimization; Finally, the three groups of 256×64×64 are averaged to fuse information from different sources and generate a representative output, which provides sufficient spatial information support for the subsequent inter-frame point-by-point mask attention module and frame reconstruction module.

3. The Transformer-based high-resolution face video processing method according to claim 2, characterized in that: In step 2, the enhanced 5-frame image information is obtained according to formula (1). In the face video super-resolution task, these 5-frame image information is integrated into the predicted intermediate frame np, and the temporal consistency method based on the inter-frame point-by-point mask attention module is adopted, as shown in formulas (2) and (3); Difference(i)=Pixel(n+i)-Pixel(n) (2) Mask:Pixel(n+i)-Pixel(n)≤0.00001 (3) Among them, formula (2) is the calculation method of pixel differences between frames, and formula (3) is the Mask identification standard; Difference represents the point-by-point difference between frames, Pixel represents the pixel of the current frame, and Mask represents the masking of unimportant pixel positions. The criteria for determining whether a pixel is important are shown in formula (3).

4. The Transformer-based high-resolution face video processing method according to claim 3, characterized in that: The calculation formula of the attention mechanism is: Among them, Q, K, and V are vectors in the attention mechanism, and their sizes are determined by the size of the input image and the size of the vector representation of each pixel; the Softmax function is a mathematical function commonly used in machine learning, d k Represents the dimension of vector K, K T Indicates the transposition of vector K, and T represents the transpose symbol of the matrix.

5. The Transformer-based high-resolution face video processing method according to claim 4, wherein in the standard attention mechanism calculation, Q is equal to the product of the input X and the parameter W, and W is subjected to a low-rank decomposition such that: The size of embed_dim is 256, and the rank is 64. In the standard attention mechanism, the parameter size corresponding to Q is the product of 256 and 256, which is 65536. After using the low-rank decomposition technology, the parameter size corresponding to Q is the product of 2, 64 and 256, which is 32768, and the number of parameters is halved. After the low-rank decomposition of the parameters corresponding to K and V, their parameter sizes are also halved respectively. By adjusting the rank, the multiple of parameter reduction can be freely controlled, and the specific reduction scale is calculated by formula (6); Scale=embed_dim÷(rank×2) (6) The attention mechanism parameter optimization method based on low-rank decomposition provides different parameter compression ratios to meet the actual application computing needs.

Citation Information

Patent Citations

  • Space-time mixed video super-resolution method based on deformable attention

    CN115861068A

  • Learable low-rank bilinear behavior perception method

    CN120071445A

  • Generative end-to-end video transmission

    CN120151533A

  • Action Recognition Method, Apparatus and Device, Storage Medium and Computer Program Product

    US20230067934A1