An Image Quality Enhancement Method Based on Facial Recognition

By performing frame processing, bone point recognition and image fusion technology on the captured video, the problem of uneven image quality is solved, and clear and coherent target images are generated, which improves the accuracy and stability of face recognition.

CN119919302BActive Publication Date: 2025-08-05BEIJING ZHONGSHITONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398142.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-05
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In practical applications, in the face of large numbers and uneven quality images, how to determine which images have the potential to effectively improve the recognition effect through enhancement technology, and thus select the right image for enhancement has become an urgent problem.

Method used

By framed the captured video, bone points are identified and abnormal frames are deleted, video frame combinations are filtered based on similarity, image quality is evaluated using PSNR, SSIM, BRISQUE, NIQE and other indicators, and target images are generated through image fusion and enhancement technologies, including denoising, defuzzing and frequency domain enhancement.

Benefits of technology

The accuracy and stability of face recognition are improved. Through effective image quality enhancement, the loss of face features caused by extreme postures is reduced, the fusion distortion is suppressed, the clarity and detail performance of the image are improved, and the recognition efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919302B_ABST
    Figure CN119919302B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of face recognition technology, and specifically discloses an image quality enhancement method based on face recognition, including the following steps: S1: Obtain video frames and number them, obtain the video frame A with the smallest number, and determine the pending number; S2: Sort the pending numbers in ascending order, determine the truncation position j in the sorting based on a preset constraint condition, and group the video frames based on the truncation position j; Remove the grouped video frames, re-number the remaining video frames and obtain new groups again until all video frames have been grouped; S3: Obtain the target image based on the image fusion technology, obtain the quality index and calculate the quality score, and determine the target image for quality enhancement based on the quality score. The present invention improves the accuracy and stability of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition, and particularly relates to an image quality enhancement method based on face recognition. Background Art

[0002] Face recognition is a biometric identification technology that identifies a person's identity based on the facial feature information of the person. It uses a collection device to collect images or video streams containing human faces, automatically detects and tracks human faces in the images, and then performs a series of related technologies for facial recognition of the detected human faces. It is usually also called portrait recognition or facial recognition.

[0003] Due to the influence of factors such as light, angle, and occlusion, the collected images may be blurred, distorted, etc., and the quality of the images is low, which affects the accuracy and speed of face recognition. For example, under direct strong light, some areas of the human face may be overexposed and lose details; or when the human face is partially occluded, key features cannot be completely extracted, resulting in difficult recognition. In addition, different shooting angles may also deform the facial contour, further increasing the recognition difficulty and reducing the recognition efficiency and accuracy. At this time, it is necessary to enhance the quality of the images used for face recognition. However, in practical applications, in the face of a large number of images with uneven quality, how to determine which images have the potential to effectively improve the recognition effect through enhancement technology, so as to select suitable images for enhancement, has become an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the present invention is to provide an image quality enhancement method based on face recognition to solve the following technical problems:

[0005] In practical applications, in the face of a large number of images with uneven quality, how to determine which images have the potential to effectively improve the recognition effect through enhancement technology, so as to select suitable images for enhancement, has become an urgent problem to be solved.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] An image quality enhancement method based on face recognition includes the following steps:

[0008] S1: Based on the video collected by the collection device when the user performs face recognition, perform frame splitting on the video to obtain video frames, number the video frames in the order of the time axis, obtain the video frame A with the smallest number, and obtain the numbers of the video frames whose similarity to the video frame A is greater than a preset value, which are recorded as pending numbers;

[0009] S2: Ascendingly sort the to-be-determined numbers, determine the truncation position j in the sorting based on a preset constraint condition, and group the video frames corresponding to the to-be-determined numbers from the first to the j-th position in the sorting and the video frame A into the same group;

[0010] Remove the grouped video frames, re-number the remaining video frames and obtain new groups again until all video frames have been grouped;

[0011] S3: Based on the image fusion technology, fuse the images in the same group to obtain a target image, obtain the quality metrics of the target image, the quality metrics include PSNR, SSIM, BRISQUE, NIQE, calculate the quality score based on the Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS), obtain the target image C with the maximum quality score, enhance the quality of the target image C, and the target image C with enhanced quality is used for face recognition.

[0012] As a further solution of the present invention: In the step S2, the constraint condition for the truncation position j is that the to-be-determined number at the i-th position in the sorting is i + 1 and the to-be-determined number at the k-th position in the sorting is not k + 1, i ∈ [1, j], k ∈ [j + 1, n], and n represents the total number of to-be-determined numbers in the sorting.

[0013] As a further solution of the present invention: In the step S1, the process of obtaining the similarity degree between the video frame and the video frame A specifically includes:

[0014] Perform gray-scale processing on the video frame A to obtain a gray-scale image, denoted as image X1;

[0015] Perform gray-scale processing on the video frame to obtain a gray-scale image, denoted as image X2;

[0016] Calculate the difference in gray-scale values of the same pixel point in the image X1 and the image X2, calculate the absolute value, and count the total absolute value. Then, the similarity degree P between the video frame and the video frame A is P = η / D, where D represents the total absolute value and η is a preset correction coefficient.

[0017] As a further solution of the present invention: If the total absolute value D = 0, it is determined that the similarity degree is greater than the preset value.

[0018] As a further solution of the present invention: In the step S3, enhancing the quality of the target image C includes: denoising, deblurring, and frequency domain enhancement.

[0019] As a further solution of the present invention: Before numbering the video frames in the step S1, the following steps are further included:

[0020] Identify the bone points in the video frame, obtain the distance between any two bone points, and Fxy represents the distance between bone point x and bone point y;

[0021] If Fxy < 0.8Fxy', then Fxy is recorded as the abnormal distance, and Fxy' represents the preset distance between the bone point x and the bone point y;

[0022] If the number of abnormal distances is greater than the preset number threshold, then the corresponding video frame is determined to be an abnormal frame, and the abnormal frame is deleted from the video frame.

[0023] As a further solution of the present invention: the bone points in the video frame are recognized based on the OpenPose algorithm.

[0024] Advantages of the present invention: Compared with the prior art:

[0025] 1) After the video is frame-divided, by detecting whether the distance between bone points is significantly less than the preset value, frames with a large side turn of the face, occlusion, or abnormal posture can be recognized, effectively reducing the loss of face features caused by extreme postures, thereby improving the effectiveness and accuracy of face recognition;

[0026] 2) By calculating the similarity between different frames and the reference frame A, video frames with a similarity greater than the preset threshold are grouped into the same group, thereby ensuring the similarity of the image in terms of texture, brightness, etc. during fusion. This can not only suppress the fusion distortion caused by excessive differences but also reduce the noise of the fused image, and finally obtain a clearer and more coherent target image;

[0027] 3) After obtaining the target image, multiple metrics such as PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity), BRISQUE, and NIQE are used in the solution for measurement, covering different dimensions such as with reference and without reference, subjective and objective, and the Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS) is introduced to comprehensively evaluate each metric, improving the objectivity of image quality assessment;

[0028] 4) The selected best target image C may still have residual noise or blurring during the fusion process. This solution improves the recognizable ability of face features by denoising and deblurring processing, making the edges of the image sharper and the details more abundant. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention will be further described below with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flowchart of an image quality enhancement method based on face recognition according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0032] Please refer to Figure 1 As shown, the present invention is an image quality enhancement method based on face recognition, including the following steps:

[0033] S1: Based on the acquisition device to collect the video of the user during face recognition, perform frame division processing on the video to obtain video frames, number the video frames in the order of the time axis, obtain the video frame A with the smallest number, and obtain the numbers of the video frames whose similarity to the video frame A is greater than a preset value, denoted as the pending numbers;

[0034] It can be understood that, first, use a camera, a smartphone camera, a monitoring device, etc. to collect the video of the user during face recognition. At this time, the system will preset the parameters of video acquisition, such as resolution (e.g., 1080p or 720p), frame rate (e.g., 25 or 30 frames per second), and color mode, to ensure that the acquired video meets the clarity requirements and is convenient for subsequent processing; use a video decoding algorithm (e.g., based on the VideoCapture module in OpenCV) to perform frame division processing on the collected video, and decompose the continuous video stream into single-frame images at a fixed time interval; after frame division processing, number the video frames according to the timestamp information or acquisition order of each frame. The numbering usually starts from 1, and the frame with the smallest number (i.e., the first frame) is the video frame A, and then the pending numbers are screened according to the similarity degree;

[0035] In a preferred embodiment of the present invention, in the step S1, the process of obtaining the similarity degree between the video frame and the video frame A specifically includes:

[0036] Perform grayscale processing on the video frame A to obtain a grayscale image, denoted as image X1;

[0037] Perform grayscale processing on the video frame to obtain a grayscale image, denoted as image X2;

[0038] Calculate the difference in grayscale values of the same pixel point in the image X1 and the image X2, and calculate the absolute value. Statistically calculate the total absolute value, then the similarity degree P between the video frame and the video frame A = η / D, D represents the total absolute value, and η is a preset correction coefficient;

[0039] It should be noted that the video frame A and other video frames to be compared are respectively subjected to grayscale processing to convert the color image into a grayscale image. The converted images are respectively denoted as image X1 and image X2. This conversion not only reduces the computational complexity but also eliminates the interference of color factors on the similarity calculation. The grayscale values of the corresponding pixel points in image X1 and image X2 are compared, the difference in the grayscale value of each pixel point is calculated and the absolute value is taken, and then the absolute differences of all pixel points are accumulated to obtain a total absolute value, which reflects the overall grayscale difference between the two frames. The greater the overall grayscale difference, the lower the similarity, and vice versa.

[0040] In a preferred case of this embodiment, if the total absolute value D = 0, it is determined that the similarity degree is greater than the preset value.

[0041] It can be understood that the total absolute value D = 0 means that the two frames are exactly the same.

[0042] In a preferred embodiment of the present invention, before numbering the video frames in step S1, the following steps are further included:

[0043] Identify the skeleton points in the video frame, and obtain the distance between any two skeleton points. Fxy represents the distance between skeleton point x and skeleton point y.

[0044] If Fxy < 0.8Fxy', then Fxy is recorded as an abnormal distance, and Fxy' represents the preset distance between skeleton point x and skeleton point y.

[0045] If the number of abnormal distances is greater than the preset number threshold, it is determined that the corresponding video frame is an abnormal frame, and the abnormal frame is deleted from the video frames.

[0046] In a preferred case of this embodiment, the skeleton points in the video frame are identified based on the OpenPose algorithm.

[0047] It should be noted that each frame of video is processed by using skeleton point detection algorithms such as OpenPose to automatically detect the positions of skeleton points of key human body parts (such as eyes, nose, ears, etc.). For each pair of key skeleton points detected (for example, left eye and right eye, etc.), the Euclidean distance formula is used to calculate their actual distance, denoted as Fxy, where x and y respectively represent two human key points; during system initialization or based on historical data, the standard distance Fxy' of each pair of skeleton points in normal front shooting is predefined. For example, in the front shooting state, the preset distance between the left eye and the right eye may be 50 pixels.

[0048] Exemplarily, in one frame, under normal circumstances, the preset distance F (left eye, right eye) between the left eye and the right eye is 50 pixels. If the distance F (left eye, right eye) between the left eye and the right eye detected in the current frame is 38 pixels, then 38 is less than 0.8 × 50 (i.e., 40 pixels). Therefore, this distance is determined to be an abnormal distance;

[0049] Perform the above comparison for all detected pairs of skeleton points within each frame, and count the number of abnormal distances. If the number of abnormal distances in a certain frame exceeds the preset number threshold (for example, exceeds 2), it is considered that there is a large abnormality in the whole frame, which may be caused by the user's posture deflection (such as tilting the head, lowering the head or partial occlusion), resulting in an abnormal shooting angle, and thus this frame is determined to be an abnormal frame; In addition to the pair of the left eye and the right eye, if there are also abnormalities in the pairs of the left ear and the right ear and the nose tip and the chin (the actual distances are all less than 0.8 times the preset value), then the cumulative number of abnormal distances reaches 3, exceeding the set threshold. At this time, this frame is determined to be an abnormal frame and will be excluded from the subsequent set of video frames and will no longer participate in subsequent processing steps such as numbering and similarity calculation;

[0050] S2: Sort the pending numbers in ascending order, determine the truncation position j in the sorting based on the preset constraint conditions, and group the video frames corresponding to the pending numbers from the first to the j-th position in the sorting and the video frame A as the same group;

[0051] Remove the grouped video frames, re-number the remaining video frames and obtain new groups again until all video frames have been grouped;

[0052] In a preferred embodiment of the present invention, in the step S2, the constraint condition for the truncation position j is: the pending number at the i-th position in the sorting is i + 1 and the pending number at the k-th position in the sorting is not k + 1, where i ∈ [1, j] and k ∈ [j + 1, n], and n represents the total number of pending numbers in the sorting.

[0053] It can be understood that, assuming the video frame numbers are 1, 2, 3, 4, 5, 6..., and the reference frame A is 1, after similarity screening, the obtained set of pending numbers is {2, 3, 5, 6, 7, 10}. After sorting this set in ascending order, its order remains 2, 3, 5, 6, 7, 10;

[0054] Traverse the sorted set of pending numbers and check whether they are consecutive starting from the first position:

[0055] If the number at the first position is 2 (i.e., 1 + 1), it meets the condition;

[0056] If the number at the second position is 3 (i.e., 2 + 1), it still meets the condition;

[0057] When the third digit is checked, if the number is not equal to 4 (i.e., 2 + 1 = 4), it indicates a break in continuity. At this time, the truncation position j is 2;

[0058] Group the reference frame A (the frame with the smallest number, usually 1) and the video frames corresponding to the pending numbers within the truncation position j (in this example, the frames numbered 2 and 3). The video frames in this group are continuous in shooting time and similar in content, which is suitable for subsequent image fusion;

[0059] After grouping, remove these video frames from the entire set of video frames to avoid reusing them in subsequent processing; renumber the remaining video frames (in this example, the remaining frame numbers are 5, 6, 7, 10, etc.) in chronological order, and then repeat the above steps; taking the remaining frames {5, 6, 7, 10} as an example, after renumbering, it may become 1, 2, 3, 4. If in the new grouping process, the set of pending numbers obtained after similarity screening is {2, 3} (some frames in the original corresponding frames 5, 6, 7), then determine the truncation position according to the continuity constraint, and finally group the new reference frame and the continuous frames into one group;

[0060] S3: Based on image fusion technology, fuse the images in the same group to obtain a target image, obtain the quality metrics of the target image, the quality metrics include PSNR, SSIM, BRISQUE, NIQE, calculate the quality score based on the Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS), obtain the target image C with the highest quality score, enhance the quality of the target image C, and the target image C with enhanced quality is used for face recognition;

[0061] In a preferred embodiment of the present invention, in step S3, enhancing the quality of the target image C includes: denoising, deblurring, and frequency domain enhancement;

[0062] It should be noted that for the continuous video frames within the same group, first preprocess the images (such as unifying the size and normalizing the brightness), and then use multi-scale fusion or weighted average method to fuse the advantages of each frame into a target image. For example, if there are five frames in the same group, through pixel-level or feature-level fusion, the best information of the clear regions of each frame can be extracted to form a target image with richer details and lower noise; a multi-scale fusion technology based on wavelet transform can be used to decompose the image into low-frequency and high-frequency parts, calculate the fusion weights for each frame respectively, and then reconstruct and synthesize the final image; or use a simple weighted average method, where the weights of each frame are adaptively adjusted according to the local contrast and texture information of the image;

[0063] After selecting the target image C, in order to further improve its clarity and detail performance and make it more suitable for face recognition, the following image enhancement processing is carried out: 1) Denoising processing: Use non-local means, bilateral filtering or deep learning denoising networks to suppress random noise in the image; for the image C with more noise, bilateral filtering can retain edge information while smoothing the noise, thus achieving the purpose of denoising; 2) Deblurring processing: For the blurring phenomenon in the image caused by motion or out-of-focus reasons, use blind deconvolution, Wiener filtering or deep learning-based deblurring algorithms for restoration. If the target image C has slight motion blurring, the blind deconvolution method can be used to estimate the blur kernel and the iterative algorithm is used to restore the image clarity; 3) Frequency domain enhancement: Use the Fourier transform to convert the image to the frequency domain, and perform enhancement processing on the high-frequency components, such as high-frequency sharpening or band-pass filtering, to highlight the image details; in the frequency domain, appropriately suppress the low-frequency components, and at the same time perform gain processing on the high-frequency components, and then perform the inverse Fourier transform to make the image edge and texture information more obvious and further improve the clarity of the key face features.

[0064] After the above fusion, quality evaluation and enhancement processing, the target image C not only reaches the optimal state in terms of overall quality, but also significant improvements have been made in various key details (such as edges, textures and contrast), which makes the face features in the image more prominent, provides a very clear and detail-rich input for the subsequent face recognition module, and thus greatly improves the recognition accuracy and robustness.

[0065] The above has described an embodiment of the present invention in detail, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A method for enhancing image quality based on face recognition, characterized in that: The following steps are involved: S1: Based on the video of the user performing face recognition captured by the acquisition device, the video is frame-processed to obtain video frames, the video frames are numbered in timeline order, the video frame A with the smallest number is obtained, and the numbers of the video frames whose similarity to the video frame A is greater than a preset value are obtained and recorded as pending numbers; S2: sorting the pending numbers in ascending order, determining a cutoff position j in the sorting based on a preset constraint condition, and grouping the video frames corresponding to the pending numbers from the first position to the jth position in the sorting and the video frame A into the same group; Remove the grouped video frames, renumber the remaining video frames and obtain new groups again until all video frames have been grouped; S3: Based on image fusion technology, images in the same group are fused to obtain a target image, and quality indicators of the target image are obtained. The quality indicators include PSNR, SSIM, BRISQUE, and NIQE. The quality score is calculated based on the superiority-inferior solution distance method, and the target image C with the largest quality score is obtained. The quality of the target image C is enhanced, and the enhanced target image C is used for face recognition.

2. The image quality enhancement method based on face recognition according to claim 1, characterized in that: In step S2, the constraint condition of the truncation position j is: the pending number of the i-th position in the sort is i+1 and the pending number of the k-th position in the sort is not k+1, i∈[1,j], k∈[j+1,n], n represents the total number of pending numbers in the sort.

3. The image quality enhancement method based on face recognition according to claim 1, characterized in that: In step S1, the process of obtaining the similarity between the video frame and the video frame A specifically includes: Performing grayscale processing on the video frame A to obtain a grayscale image, which is recorded as image X1; Perform grayscale processing on the video frame to obtain a grayscale image, which is recorded as image X2; Calculate the grayscale value difference of the same pixel point in the image X1 and the image X2, calculate the absolute value, and count the total absolute value. Then the similarity between the video frame and the video frame A is P=η / D, where D represents the total absolute value and η is a preset correction coefficient.

4. The image quality enhancement method based on face recognition according to claim 3, characterized in that: If the total absolute value D=0, it is determined that the similarity is greater than the preset value.

5. The image quality enhancement method based on face recognition according to claim 1, characterized in that: In step S3, enhancing the quality of the target image C includes: denoising, deblurring and frequency domain enhancement.

6. The image quality enhancement method based on face recognition according to claim 1, characterized in that: In the step S1, before numbering the video frames, the following steps are also included: Identify the skeleton points in the video frame and obtain the distance between any two skeleton points, where Fxy represents the distance between skeleton point x and skeleton point y; If Fxy < 0.8Fxy', then Fxy is recorded as the abnormal distance, and Fxy' represents the preset distance between bone point x and bone point y; If the number of abnormal distances is greater than a preset number threshold, the corresponding video frame is determined to be an abnormal frame, and the abnormal frame is deleted from the video frame.

7. The image quality enhancement method based on face recognition according to claim 6, characterized in that: Skeleton points in the video frame are identified based on the OpenPose algorithm.

Citation Information

Patent Citations

  • Video face identifying method

    CN104008370A

  • Face recognition processing method and device

    CN115880753A