Face super-resolution recognition method and system based on video stream
Through a 3D convolutional neural network and attention mechanism based on video stream, combined with deep separable convolution and multi-stage upsampling strategy, the problem of low recognition accuracy of low-resolution face images is solved, and efficient and high-precision face super-resolution and recognition are achieved.
Patent Information
- Application Number
- CN202510725701.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, low-resolution facial image recognition has low accuracy and poor computational efficiency. Traditional methods fail to fully utilize the spatiotemporal information of adjacent frames in the video stream, and the separation of super-resolution and recognition processes leads to error accumulation.
Through the 3D convolutional neural network based on video stream combined with channel and spatial attention mechanism, spatiotemporal feature extraction and fusion are performed, and an end-to-end joint network is adopted for face super-resolution and feature extraction. Deep separable convolution and multi-stage upsampling strategy are used, combined with super-resolution and recognition loss function optimization.
It achieves high-precision super-resolution reconstruction and efficient recognition, restores the details of low-resolution facial images, and improves the accuracy and computational efficiency of face recognition.
Smart Images

Figure CN120673456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and artificial intelligence, and in particular to a method and system for super-resolution face recognition based on video streams. Background Art
[0002] As an important branch of biometric identification, facial recognition technology has been widely used in security monitoring, identity verification, smart access control, and other fields. However, in real-world applications, facial images often suffer from low resolution and blur due to factors such as the resolution limitations of monitoring equipment, long shooting distances, and poor lighting conditions. This seriously affects facial recognition accuracy.
[0003] Traditional face super-resolution methods primarily process single-frame, low-resolution facial images, increasing image resolution through interpolation and reconstruction techniques. While these methods can improve image quality to a certain extent, due to the limited information contained in a single frame, it is difficult to recover detailed facial features, resulting in suboptimal super-resolution results and, in turn, impacting face recognition accuracy.
[0004] In recent years, with the development of deep learning technology, super-resolution methods based on convolutional neural networks have made significant progress. By learning from a large number of low-resolution and high-resolution image pairs, these methods can automatically extract image features and reconstruct them, effectively improving super-resolution results. However, most existing deep learning methods still process single frames and fail to fully utilize the spatiotemporal information between adjacent frames in the video stream.
[0005] Furthermore, traditional face super-resolution and recognition methods typically operate separately, first super-resolutioning a low-resolution face image and then feeding the processed image into a face recognition model for recognition. This separate approach not only increases computational complexity but also further reduces face recognition accuracy due to potential errors introduced during the super-resolution process. Summary of the Invention
[0006] The purpose of the present invention is to provide a face super-resolution recognition method and system based on video stream. By fully mining the spatiotemporal information of consecutive frames in the video stream and combining it with deep learning technology, high-precision face super-resolution reconstruction and accurate face recognition are achieved, thereby solving the problems of low recognition accuracy and poor computational efficiency of low-resolution face images in the existing technology.
[0007] The technical solutions of the present invention are as follows:
[0008] A face super-resolution recognition method based on video stream, comprising the following steps:
[0009] Obtain a continuous video frame sequence, pre-process each frame and detect the face area, and extract the face ROI;
[0010] Extract and fuse spatiotemporal features of continuous multi-frame face ROIs;
[0011] The fused feature map is then used to perform face super-resolution and feature extraction simultaneously through an end-to-end joint network;
[0012] Match the extracted feature vector with the database template and output the recognition result.
[0013] Furthermore, the preprocessing step includes: gray-scaling the video frame, enhancing the image contrast by using a histogram equalization method, removing image noise by using a Gaussian filter, normalizing the image by an affine transformation, and adjusting the image size to a uniform size.
[0014] Furthermore, the spatiotemporal feature extraction and fusion of the face ROIs of consecutive multiple frames are specifically as follows:
[0015] The facial ROIs of multiple consecutive frames are input into the 3D convolutional neural network, and features are extracted in the time and space dimensions through the 3D convolution kernel; the channel attention mechanism is used to perform average pooling and maximum pooling operations on the input feature map respectively, and the results are input into the multi-layer perceptron respectively, and then the results output by the multi-layer perceptron are added and passed through the activation function to obtain the channel attention weight; through the spatial attention mechanism, the input feature map is average pooling and maximum pooling operations in the channel dimension, the two pooling results are spliced, and then the spatial attention weight is generated through the convolution operation; the channel attention weight and spatial attention weight are multiplied with the feature map extracted by the 3D convolutional neural network to realize the fusion of spatiotemporal features.
[0016] Furthermore, the training loss function of the joint network includes a weighted combination of super-resolution loss and recognition loss, wherein the super-resolution loss adopts perceptual loss and adversarial loss, and the recognition loss adopts an angle-based ArcFace loss function.
[0017] Furthermore, the fused feature map is simultaneously subjected to face super-resolution and feature extraction through an end-to-end joint network. Specifically, the fused feature map is input into the shared feature extraction layer to extract shared features; the shared features are input into the super-resolution branch and the feature extraction branch respectively; the super-resolution branch converts the shared features into a high-resolution face image through residual blocks and sub-pixel convolution layers; the feature extraction branch maps the shared features into feature vectors through multiple fully connected layers.
[0018] Furthermore, the matching of the extracted feature vector with the database template is specifically as follows:
[0019] The extracted feature vectors are L2 normalized, the cosine similarity of the normalized feature vectors is calculated, a threshold is set for matching judgment, and the corresponding identity information is output when the similarity is greater than the set threshold.
[0020] Furthermore, the residual block of the super-resolution branch uses depthwise separable convolution instead of standard convolution to reduce the number of parameters and improve computational efficiency; the sub-pixel convolution layer adopts a multi-stage upsampling strategy to improve image resolution while retaining detail information.
[0021] The present invention also provides a face super-resolution recognition system based on video stream, the system comprising:
[0022] The video acquisition module is used to obtain a continuous video frame sequence, pre-process each frame, detect the face area, and extract the face ROI;
[0023] The spatiotemporal feature fusion module is used to extract and fuse spatiotemporal features of multiple consecutive frames of face ROI;
[0024] The feature extraction module is used to perform face super-resolution and feature extraction on the fused feature map through an end-to-end joint network;
[0025] The result matching module is used to match the extracted feature vector with the database template and output the recognition result
[0026] Compared with the prior art, the present invention has the following advantages:
[0027] High-precision super-resolution reconstruction: By combining 3D convolutional neural networks with channel attention and spatial attention mechanisms, the spatiotemporal information of consecutive frames in the video stream is fully exploited, which can effectively restore the details of low-resolution facial images, achieve high-quality super-resolution reconstruction, and provide richer feature information for subsequent recognition.
[0028] Efficient recognition: An end-to-end joint network is used to simultaneously perform face super-resolution and feature extraction, reducing intermediate processing steps and improving computational efficiency. The super-resolution branch uses depthwise separable convolution and multi-stage upsampling strategies to reduce computational complexity while ensuring reconstruction quality.
[0029] High Accuracy: The training loss function of the joint network combines super-resolution loss and recognition loss. By optimizing the perceptual loss, adversarial loss, and ArcFace loss functions, the reconstructed high-resolution face images are closer to the real images, and the extracted feature vectors have stronger discrimination, thereby significantly improving the accuracy of face recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings illustrate various embodiments generally by way of example and not limitation, and together with the description and claims, serve to explain embodiments of the invention. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive of the embodiments of the present apparatus or method.
[0031] Figure 1 A flow chart of the method of the present invention is shown. DETAILED DESCRIPTION
[0032] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0033] like Figure 1 As shown, the present invention provides a face super-resolution recognition method based on video stream, comprising:
[0034] Video frame acquisition and preprocessing:
[0035] A continuous sequence of video frames is acquired through a video acquisition device (such as a camera). Each video frame is first grayscaled to convert the color image into a grayscale image, reducing the data dimension. Histogram equalization is then used to expand the image's grayscale dynamic range, enhance image contrast, and make image details clearer. Gaussian filtering is then used to smooth the image and remove noise, preventing it from interfering with subsequent processing. The image is then normalized using an affine transformation to adjust the facial region in the image to a standard position and pose. Finally, the image is resized to a uniform size, such as 128×128 pixels, for subsequent processing. Existing, mature face detection algorithms (such as MTCNN) are used to detect the facial region in the preprocessed image and extract the facial ROI.
[0036] Spatiotemporal feature extraction and fusion:
[0037] Select N consecutive frames of face ROI (usually based on experiments and actual results, N can be 5-10, here N = 5 is taken as an example) and input them into the 3D convolutional neural network.
[0038] The 3D convolutional neural network consists of five convolutional layers, all with kernel sizes of (3,3,3), stride size of 1, padding of 1, and the number of channels set to 32, 64, 128, 256, and 512, respectively. Compared to traditional 2D convolutional neural networks, the 3D convolutional kernels of 3D convolutional neural networks can simultaneously extract features from multiple input frames in both the temporal and spatial dimensions (the length and width of the image). In the temporal dimension, they can capture facial motion between consecutive frames, such as changes in facial expression (smiling, frowning, etc.), head rotation, and blinking. In the spatial dimension, they can extract static structural features of the face, such as the shape, positional relationship, and facial contours of the facial features. This combined extraction of temporal and spatial features enables the network to more comprehensively describe facial features. For example, in surveillance videos, when a person walks or speaks, the 3D convolutional neural network can accurately capture these dynamic changes through temporal feature extraction. Combined with spatial features, it effectively extracts facial features in complex scenes.
[0039] Channel Attention Module: The MLP hidden layer dimension is set to 1 / 16 of the number of channels, using the ReLU activation function, and the output layer uses a Sigmoid activation. In the feature maps output by a 3D convolutional neural network, different channels may contain information of varying types and importance. The channel attention module automatically learns the importance of each channel. By performing average pooling and max pooling on the feature maps, it captures global information from different perspectives. Average pooling reflects the overall mean of the feature map, while max pooling highlights significant features within the feature map. The two pooling results are fed into a multi-layer perceptron (MLP). After the MLP's nonlinear transformation, the two outputs are added and processed through an activation function to generate channel attention weights. These weights are used to adjust the importance of different channel features, enabling the network to prioritize key information channels and suppress irrelevant or secondary information channels. For example, in facial features, information related to key areas such as the eyes and nose may be more important. Channel attention weights increase the weights of these channels accordingly, allowing the network to better extract and utilize this key information.
[0040] Spatial attention module: The convolution kernel size is 7×7, the number of channels is 1, and Sigmoid activation is used.
[0041] In the channel attention mechanism, average pooling and max pooling are performed on the feature maps output by the 3D convolutional neural network, respectively. Average pooling calculates the average value within a local region of the feature map, while max pooling selects the maximum value within that region. These two pooling methods allow for global information about the feature map to be captured from different perspectives. The two pooling results are then fed into a multi-layer perceptron (MLP), a feedforward neural network consisting of an input layer, several hidden layers, and an output layer. After the MLP's nonlinear transformation, the two outputs are added together and processed through an activation function to obtain the channel attention weights. These weights reflect the importance of different channel features within the overall feature map. By adjusting the channel weights, attention can be focused on key information channels while suppressing irrelevant or secondary information channels.
[0042] In the spatial attention mechanism, average pooling and max pooling are performed on the feature map in the channel dimension, and the two pooling results are concatenated to form a new feature representation. Then, a convolution operation is performed, in which the convolution kernel slides over this concatenated feature map, generating spatial attention weights through the convolution operation. These weights are used to highlight important spatial regions in the feature map, allowing the network to focus more on key facial features such as the eyes, nose, and mouth, which are highly recognizable. Finally, the channel attention weights and spatial attention weights are multiplied with the feature map extracted by the 3D convolutional neural network to achieve effective fusion of spatiotemporal features, resulting in a feature map containing rich spatiotemporal information, providing more valuable input for subsequent face super-resolution and feature extraction.
[0043] Face super-resolution and feature extraction:
[0044] The fused feature map is input into the end-to-end joint network, which is divided into a shared feature extraction layer, a super-resolution branch, and a feature extraction branch. The specific process is as follows:
[0045] Shared feature extraction layer: It consists of multiple convolutional layers and pooling layers. The convolutional layer uses convolution kernels to slide on the feature map to perform convolution operations. By setting different convolution kernel sizes and step sizes, features of different scales and directions can be extracted. For example, smaller convolution kernels (such as 3×3) can extract local detail features, while larger convolution kernels (such as 5×5 or 7×7) are more suitable for extracting global structural features. The pooling layer reduces the size of the feature map and reduces the amount of calculation through downsampling operations (such as maximum pooling or average pooling), while retaining the main feature information. Maximum pooling can highlight the significant features in the feature map, while average pooling focuses more on preserving the overall feature distribution. After processing by the shared feature extraction layer, shared feature representations are obtained. These features contain the basic feature information of the facial image and provide a basis for subsequent super-resolution reconstruction and feature extraction.
[0046] Super-resolution branch: Super-resolution reconstruction of face images is achieved using residual blocks and sub-pixel convolutional layers.
[0047] Residual block: Using depthwise separable convolution, the standard convolution is decomposed into depthwise convolution and pointwise convolution. Each residual block contains two convolutional layers, with the ReLU activation function used in the middle, and finally the input and output are added through a jump connection. This decomposition method of depthwise separable convolution significantly reduces the number of parameters and computational complexity compared to standard convolution. At the same time, the skip connection mechanism of the residual block plays an important role. It can directly add the input to the output after the convolution operation, avoiding the gradient vanishing problem during network training. This enables the network to learn deeper features, improving the network's training stability and reconstruction effect. In practical applications, even if the network has a deep number of layers, the residual block can ensure effective network training, thereby achieving high-quality super-resolution reconstruction.
[0048] Sub-pixel convolutional layer: This layer uses a multi-stage upsampling strategy, such as three-stage upsampling, with each stage upsampling by a factor of 2. The image resolution is gradually increased through upsampling at lower magnifications, and then the feature map is converted into a high-resolution image through pixel reassembly. For example, a 128×128 low-resolution image is reconstructed into a 512×512 high-resolution image after three stages of upsampling. This multi-stage upsampling approach gradually restores image details while maintaining image quality, avoiding the blurring and distortion that can occur with a single upsampling step. In the reconstructed high-resolution image, facial details such as hair and skin texture are well restored, providing a clearer image for subsequent face recognition.
[0049] Feature extraction branch: Multiple fully connected layers are used to map shared features into a feature space, generating feature vectors for face recognition. Each neuron in a fully connected layer is connected to all neurons in the previous layer, enabling global nonlinear transformations of input features. To improve the generalization and discrimination of the feature vectors, batch normalization and dropout layers are added after the fully connected layers.
[0050] Batch Normalization layer: By normalizing the input data, it accelerates network convergence and reduces training time. During deep learning network training, changes in data distribution can lead to instability. Batch Normalization normalizes the data within each mini-batch, making the data distribution more stable and improving the network's generalization capabilities. For example, facial images captured under different lighting conditions or shooting angles can be fed into the network with a more consistent distribution after batch normalization, improving the network's adaptability to different scenarios.
[0051] Dropout layer: During training, the outputs of some neurons are randomly set to 0 to prevent overfitting and enhance model robustness. Overfitting is a common problem in deep learning. When a network performs well on training data but degrades on test data, overfitting may occur. The dropout layer randomly discards the outputs of some neurons, forcing the network to learn more robust and generalizable features, enabling it to perform better when faced with new data. The dropout probability is typically set to 0.5, a suitable parameter value verified in numerous experiments, which strikes a good balance between avoiding overfitting and ensuring network performance.
[0052] Feature matching and result output:
[0053] The extracted feature vectors are matched with template feature vectors pre-stored in the database. First, the extracted feature vectors are L2 normalized to ensure that all feature vectors have the same modulus, eliminating the impact of vector length differences on similarity calculations. The cosine similarity between the normalized feature vectors and the feature vectors of each template in the database is then calculated. The formula for cosine similarity is Cosine Similarity(A,B) = (A·B) / (||A|| × ||B||), where A and B are two feature vectors, A·B represents the vector dot product, and ||A|| and ||B|| represent the moduli of vectors A and B, respectively. Cosine similarity reflects the degree of directional similarity between two vectors and ranges from [-1 to 1]. Values closer to 1 indicate greater similarity between the two vectors. Set an appropriate similarity threshold (such as 0.8, which can be adjusted according to the actual application scenario and the characteristics of the data set). When the calculated similarity is greater than the set threshold, the match is considered successful and the identity information of the corresponding template is output; if all similarities are lower than the set threshold, it is determined to be unmatched and an unknown identity is output.
[0054] The above is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A face super-resolution recognition method based on video stream, characterized in that: The following steps are involved: Obtain a continuous video frame sequence, pre-process each frame and detect the face area, and extract the face ROI; Extract and fuse spatiotemporal features of continuous multi-frame face ROIs; The fused feature map is then used to perform face super-resolution and feature extraction simultaneously through an end-to-end joint network; Match the extracted feature vector with the database template and output the recognition result.
2. The face super-resolution recognition method based on video stream according to claim 1, characterized in that The preprocessing steps include: gray-scaling the video frames, enhancing the image contrast by using a histogram equalization method, removing image noise by using a Gaussian filter, normalizing the image by an affine transformation, and adjusting the image size to a uniform size.
3. The face super-resolution recognition method based on video stream according to claim 1, characterized in that The spatiotemporal feature extraction and fusion of continuous multi-frame face ROIs are specifically as follows: The facial ROI of multiple consecutive frames is input into the 3D convolutional neural network, and features are extracted in the time and space dimensions through the 3D convolution kernel. The channel attention mechanism is used to perform average pooling and maximum pooling operations on the input feature map, and the obtained results are respectively input into the multi-layer perceptron. The output results of the multi-layer perceptron are then added and passed through the activation function to obtain the channel attention weight. Through the spatial attention mechanism, the input feature map is average pooled and max pooled in the channel dimension, the two pooling results are spliced, and then the spatial attention weight is generated through the convolution operation; the channel attention weight and the spatial attention weight are multiplied with the feature map extracted by the 3D convolutional neural network to realize the fusion of spatiotemporal features.
4. The face super-resolution recognition method based on video stream according to claim 1, characterized in that The training loss function of the joint network includes a weighted combination of super-resolution loss and recognition loss, where the super-resolution loss adopts perceptual loss and adversarial loss, and the recognition loss adopts an angle-based ArcFace loss function.
5. The face super-resolution recognition method based on video stream according to claim 1, characterized in that The fused feature map is used to perform face super-resolution and feature extraction simultaneously through an end-to-end joint network. Specifically, the fused feature map is input into the shared feature extraction layer to extract shared features; the shared features are input into the super-resolution branch and the feature extraction branch respectively; the super-resolution branch converts the shared features into a high-resolution face image through residual blocks and sub-pixel convolution layers; the feature extraction branch maps the shared features into feature vectors through multiple fully connected layers.
6. The method for face super-resolution recognition based on video stream according to claim 1, characterized in that: The matching of the extracted feature vector with the database template is specifically as follows: The extracted feature vectors are L2 normalized, the cosine similarity of the normalized feature vectors is calculated, a threshold is set for matching judgment, and the corresponding identity information is output when the similarity is greater than the set threshold.
7. The method for face super-resolution recognition based on video stream according to claim 5, characterized in that: The residual block of the super-resolution branch uses depthwise separable convolution instead of standard convolution to reduce the number of parameters and improve computational efficiency; the sub-pixel convolution layer adopts a multi-stage upsampling strategy to improve image resolution while retaining detail information.
8. A face super-resolution recognition system based on video stream, characterized in that: The system comprises: The video acquisition module is used to obtain a continuous video frame sequence, pre-process each frame, detect the face area, and extract the face ROI; The spatiotemporal feature fusion module is used to extract and fuse spatiotemporal features of multiple consecutive frames of face ROI; The feature extraction module is used to perform face super-resolution and feature extraction on the fused feature map through an end-to-end joint network; The result matching module is used to match the extracted feature vector with the database template and output the recognition result.
Citation Information
Cited By
Fast moving face recognition method and system based on multi-frame image enhancement
CN121214529A