Method and device for improving face quality of picture figure

By constructing a low-rank fine-tuning model to fine-tune the diffusion model, the problem of insufficient improvement of facial image details, light and shadow effects and expression nature in the prior art is solved, and high-quality facial feature image generation is achieved.

CN120070215APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108254.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has limitations in improving the details, light and shadow effects and naturalness of facial images, resulting in the generated images requiring tedious repairs by artists.

Method used

By constructing a low-rank fine-tuning model, the trained low-rank fine-tuning model is used to fine-tune the diffusion model to generate enhanced facial feature images and improve facial details, light and shadow effects and expression nature.

Benefits of technology

It has achieved a significant improvement in facial details, light and shadow effects and naturalness of expressions, and improved the quality of generated images, allowing users to directly use the enhanced images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070215A_ABST
    Figure CN120070215A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for improving the face quality of a picture character, and the method comprises the steps: screening out a needed picture from a face image library, and taking the needed picture as training image data; processing the training image data to obtain an image with a set resolution as training image processing data; constructing a low-rank fine tuning model, inputting the training image processing data into the initial low-rank fine tuning model, and performing model training to obtain a trained low-rank fine tuning model; performing face feature extraction on the image with the face to obtain a face feature vector; and inputting the face feature vector into the diffusion model, and performing fine tuning on the diffusion model through the trained low-rank fine tuning model to generate an enhanced face feature image, so that the facial details, the light and shadow effect and the expression natural effect are greatly improved, and the quality of the generated image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and particularly to a method and device for improving the facial quality of human figures in pictures. Background Art

[0002] In the field of facial image processing, existing methods have limitations in improving facial quality. For example, in dealing with facial details, lighting effects, and naturalness of expressions, etc., the ideal effects are often not achieved; as a result, the pictures generated by the diffusion model still require a series of cumbersome repairs by graphic designers, which is not convenient for users. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for improving the facial quality of human figures in pictures, so as to greatly improve facial details, lighting effects, and natural expression effects, and improve the quality of the generated pictures.

[0004] In a first aspect, the present invention provides a method for improving the facial quality of human figures in pictures, including the following steps:

[0005] Step 1: Screen out the required pictures from the face picture library as training image data;

[0006] Step 2: Process the training image data to obtain an image with a set resolution as training image processing data;

[0007] Step 3: Construct a low-rank fine-tuning model, input the training image processing data into the initial low-rank fine-tuning model for model training to obtain a trained low-rank fine-tuning model;

[0008] Step 4: Extract face features from the image with a face, obtain a face feature vector; input the face feature vector into the diffusion model, and fine-tune the diffusion model through the trained low-rank fine-tuning model to generate an enhanced facial feature image.

[0009] In a second aspect, the present invention provides a device for improving the facial quality of human figures in pictures, including:

[0010] A data screening module that screens out the required pictures from the face picture library as training image data;

[0011] A data processing module that processes the training image data to obtain an image with a set resolution as training image processing data;

[0012] A model construction and training module that constructs a low-rank fine-tuning model, inputs the training image processing data into the initial low-rank fine-tuning model for model training to obtain a trained low-rank fine-tuning model;

[0013] A feature enhancement module extracts facial features from an image with a face to obtain a facial feature vector; inputs the facial feature vector into a diffusion model, and fine-tunes the diffusion model through a trained low-rank fine-tuning model to generate an enhanced facial feature image.

[0014] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0015] Through the generation and training of the low-rank fine-tuning model, the present invention obtains a model that meets the requirements and uses it in the diffusion model, enabling the diffusion model to enhance the features of the face in the generated human image, greatly improving the facial details, lighting effects, and natural expression effects, and allowing users to directly use it for commercial display.

[0016] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings

[0017] The present invention will be further described below with reference to the drawings in conjunction with embodiments.

[0018] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0019] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0020] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0021] 1. Achieve high-definition texture enhancement of the face;

[0022] 2. Achieve the mixing of multiple low-rank fine-tuning models;

[0023] Prepare the training dataset with the following steps: First, execute the detection function using the yolov5face algorithm to identify and filter out the frontal face parts. For eligible face images, execute the calculate_face_area() function to ensure that the face area occupies 60%-80% of the central area of the image. Use the calculate_rotation() function to check if the side rotation angle is within ±5°. Then, perform color histogram analysis through the analyze_histogram() function to eliminate abnormally exposed samples, ensuring that the standard deviation of the RGB three channels of the image is less than 45 and the brightness is concentrated in the range of 0.4-0.8. Finally, use select_high_quality_images() to obtain 10,000 high-quality training images that meet the requirements.

[0024] Before training, perform preprocessing operations according to the following method. Process all images to a resolution of 512×512 using the cv2.resize() function and the INTER_AREA interpolation method. First, scale the image proportionally through the scale_image() function so that the shorter side reaches 1024 pixels. Then, use the crop_image_center() function to crop the image centered on the face center point. Normalize the pixel values from [0,255] to [-1,1] through the normalize_pixels() function, and perform random horizontal flipping and color jittering using the random_flip() and random_color_jitter() functions to enhance data diversity. Finally, set batch_size = 32 for training.

[0025] Construct a low-rank fine-tuning framework: First, decompose the weight matrix W0 of UNet into the form of W0 + BA, where B ∈ R^(d×r), A ∈ R^(r×d), and r is set to 8. Design three independent fine-tuning branches, each branch containing an independent pair of BA matrices, which focus on texture enhancement, light and shadow optimization, and expression naturalness improvement respectively.

[0026] Implement the forward propagation process of the branch. Each branch receives the input image x and performs feature extraction and transformation in the form of W0 + B i A i where i represents the branch index. The synthesis network generates enhanced images through upsampling operations using the transformed features.

[0027] The training process adopts an alternating optimization strategy. First, fix the W0 matrix and train the BA matrix; then, adjust the weights of each branch according to the evaluation results of the loss function. Use the Adam optimizer with a learning rate set to 0.001 and momentum parameters beta1 = 0.0, beta2 = 0.99.

[0028] Stabilize the training process using gradient penalty and batch normalization techniques. All convolutional layers use 3×3 or 1×1 convolutional kernels, and the activation function is LeakyReLU(α = 0.2). Regularly check the gradient norm to prevent gradient vanishing or explosion.

[0029] Use the pre-trained FaceNet model for feature extraction to obtain 512-dimensional face feature vectors. These feature vectors encode key information such as facial contours and facial feature structures, serving as guiding information for low-rank fine-tuning.

[0030] In the inference stage, the input image is sequentially processed through three fine-tuning branches. Use the DDIM sampler for feature enhancement, set the sampling steps to 50, and the time step to 0.02. At the same time, introduce attention guidance (weight 1.5) to enhance the attention of key regions.

[0031] Enhance facial features through an iterative optimization process. The model focuses on three aspects: texture details, lighting effects, and expression naturalness. Finally, upscale the output to a resolution of 2048×2048 through a 2x upsampling network.

[0032] The entire solution has achieved a significant improvement in facial details, lighting effects, and expression naturalness on the test set.

[0033] Example 1

[0034] As Figure 1 shown, this example provides a method for improving the facial quality of picture characters, including the following steps:

[0035] Step 1: Screen out the required pictures from the face image library as training image data;

[0036] Step 2: Process the training image data to obtain an image with a set resolution as the processed training image data;

[0037] Step 3: Build a low-rank fine-tuning model, input the processed training image data into the initial low-rank fine-tuning model for model training to obtain the trained low-rank fine-tuning model;

[0038] Step 4: Extract face features from the image with a face, obtain face feature vectors; input the face feature vectors into the diffusion model, and fine-tune the diffusion model through the trained low-rank fine-tuning model to generate an image with enhanced facial features.

[0039] In this embodiment, preferably, step 1 is specifically as follows: First, use the yolov5face algorithm to execute the detection function to identify and screen out the frontal view face part; for the screened face images, use the calculate_face_area function to ensure that the face area occupies 60%-80% of the central area of the image; use the calculate_rotation function to check whether the side rotation angle is within ±5°; then, perform color histogram analysis through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the RGB three channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8 to obtain preprocessed image data; use the select_high_quality_images function to screen the preprocessed data. The select_high_quality_images function first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image to a grayscale image, then uses cv2.Laplacian(gray,cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to the uint8 type. Finally, calculate the average value of all results to obtain the sharpness score. Set the threshold to 6, and screen out 10,000 images with a sharpness score greater than 6 from the preprocessed image data to obtain training image data.

[0040] In this embodiment, preferably, step 2 is specifically as follows: Process all the images in the training image data using the cv2.resize function and the INTER_AREA interpolation method, and process each image to a resolution of 512×512; use the random_flip function and the random_color_jitter function to perform random horizontal flipping and color jitter on the 512×512 resolution images, and then normalize the pixel values from [0,255] to [-1,1] through the normalize_pixels function to obtain training image processing data.

[0041] In this embodiment, preferably, step 3 is specifically as follows: Construct a low-rank fine-tuning model: First, decompose the weight matrix W0 of UNet into the form of W0+BA, where B∈R^(d×r), A∈R^(r×d), r is set to 8, and d is the dimension of the input feature; design three independent fine-tuning branches, each fine-tuning branch contains an independent BA matrix pair, and the three independent fine-tuning branches are respectively used to focus on texture enhancement improvement, light and shadow optimization improvement, and expression naturalness improvement; each branch receives the input image x, and passes through W0+B i A iFeature extraction and transformation are performed in the form of, where i represents the branch index; set batch_size = 32, input the training image processing data into the initial low-rank fine-tuning model, perform model training, and obtain the trained low-rank fine-tuning model;

[0042] The training adopts an alternating optimization strategy; fix the W0 matrix and train the BA matrix; then, adjust the weights of each branch according to the evaluation result of the loss function, and stop training when the loss function reaches the set threshold;

[0043] Use gradient penalty and batch normalization techniques to stabilize the training process; and use 3×3 or 1×1 convolution kernels for all convolutional layers, and use LeakyReLU as the activation function with α = 0.2; set an interval time to check the gradient norm.

[0044] In this embodiment, preferably, step 4 is specifically: input the image with a face into the pre-trained FaceNet model for face feature extraction to obtain a 512-dimensional face feature vector; input the 512-dimensional face feature vector into the diffusion model, and fine-tune the diffusion model through the trained low-rank fine-tuning model to generate an intermediate image with a resolution of 1024×1024, and the intermediate image is output and enlarged to a resolution of 2048×2048 through a 2x upsampling network to obtain an enhanced facial feature image.

[0045] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0046] Embodiment 2

[0047] As Figure 2 shown, in this embodiment, an apparatus for improving the facial quality of picture characters is provided, including:

[0048] A data screening module that screens out the required pictures from the face image library as training image data;

[0049] A data processing module that processes the training image data to obtain an image with a set resolution as training image processing data;

[0050] A training construction module that constructs a low-rank fine-tuning model, inputs the training image processing data into the initial low-rank fine-tuning model, performs model training, and obtains the trained low-rank fine-tuning model;

[0051] A feature enhancement module that performs face feature extraction on the image with a face to obtain a face feature vector; inputs the face feature vector into the diffusion model, and fine-tunes the diffusion model through the trained low-rank fine-tuning model to generate an enhanced facial feature image.

[0052] In this embodiment, preferably, the screening data module is specifically as follows: First, use the yolov5face algorithm to execute the detection function to identify and screen out the frontal view face part; for the screened face images, use the calculate_face_area function to ensure that the face area occupies 60%-80% of the central area of the image; use the calculate_rotation function to check whether the side rotation angle is within ±5°; then, perform color histogram analysis through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the RGB three channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8, to obtain preprocessed image data; use the select_high_quality_images function to screen the preprocessed data. The select_high_quality_images function first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image to a grayscale image, and then uses cv2.Laplacian(gray, cv2.CV_64F) to calculate the Laplacian operator, take the absolute value of the result, and convert it to the uint8 type. Finally, calculate the mean value of all results to obtain the sharpness score. Set the threshold to 6, and screen out 10,000 images with a sharpness score greater than 6 from the preprocessed image data to obtain training image data.

[0053] In this embodiment, preferably, the processing data module is specifically as follows: Process all the images in the training image data using the cv2.resize function and the INTER_AREA interpolation method, and process each image to a resolution of 512×512; perform random horizontal flipping and color jitter on the 512×512 resolution images using the random_flip function and the random_color_jitter function, and then normalize the pixel values from [0, 255] to [-1, 1] through the normalize_pixels function to obtain training image processing data.

[0054] In this embodiment, preferably, the building training module is specifically as follows: Build a low-rank fine-tuning model: First, decompose the weight matrix W0 of UNet into the form of W0+BA, where B∈R^(d×r), A∈R^(r×d), r is set to 8, and d is the dimension of the input feature; design three independent fine-tuning branches, each fine-tuning branch contains an independent pair of BA matrices, and the three independent fine-tuning branches are respectively used to focus on texture enhancement improvement, light and shadow optimization improvement, and expression naturalness improvement; each branch receives the input image x, through W0+B i A iFeature extraction and transformation are performed in the form of, where i represents the branch index; set batch_size = 32, input the training image processing data into the initial low-rank fine-tuning model, and perform model training to obtain the trained low-rank fine-tuning model;

[0055] The training adopts an alternating optimization strategy; fix the W0 matrix and train the BA matrix; subsequently, adjust the weights of each branch according to the evaluation result of the loss function, and stop training when the loss function reaches the set threshold;

[0056] Use gradient penalty and batch normalization techniques to stabilize the training process; and use 3×3 or 1×1 convolutional kernels for all convolutional layers, and use LeakyReLU as the activation function with α = 0.2; set an interval time to check the gradient norm.

[0057] In this embodiment, preferably, the feature enhancement module is specifically: input the image with a human face into the pre-trained FaceNet model for face feature extraction to obtain a 512-dimensional face feature vector; input the 512-dimensional face feature vector into the diffusion model, and fine-tune the diffusion model through the trained low-rank fine-tuning model to generate an intermediate image with a resolution of 1024×1024, and the intermediate image is output and enlarged to a resolution of 2048×2048 through a 2x upsampling network to obtain an enhanced facial feature image.

[0058] Since the device introduced in the second embodiment of the present invention is the device used to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be repeated here. Any device used in the method of the first embodiment of the present invention belongs to the scope to be protected by the present invention.

[0059] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0060] This embodiment can be used to enhance the facial features in the face pictures generated by the user using the AI function, which can enhance the facial features, greatly improve the facial details, lighting effects, and natural expression effects, so that the user can directly use the enhanced pictures without further graphic design processing.

[0061] Although the specific implementation manners of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope protected by the claims of the present invention.

Claims

1. A method for improving the facial quality of a person in an image, characterized by: The steps include: Step 1: Filter the required pictures from the face library as training image data; Step 2: Process the training image data to obtain an image with a set resolution as training image processing data; Step 3: construct a low-rank fine-tuning model, input the training image processing data into the initial low-rank fine-tuning model, perform model training, and obtain a trained low-rank fine-tuning model; Step 4: extract facial features from the image with a face to obtain a facial feature vector; input the facial feature vector into the diffusion model, and fine-tune the diffusion model through the trained low-rank fine-tuning model to generate an enhanced facial feature image.

2. The method for improving the facial quality of a person in an image according to claim 1, characterized in that: The step 1 is specifically as follows: first, the detection function is executed using the yolov5face algorithm to identify and screen out the face portion of the front view; the screened face image is used to use the calculate_face_area function to ensure that the face area occupies 60%-80% of the center area of ​​the image; the calculate_rotation function is used to check whether the side rotation angle is within ±5°; then, the color histogram analysis is performed through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the three RGB channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8, and obtain the preprocessed image data; The preprocessed image data is screened using the select_high_quality_images function, which first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image into a grayscale image, and then uses cv2.Laplacian(gray,cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to uint8 type, and finally calculates the average of all results to obtain a clarity score, sets the threshold to 6, and screens out 10,000 images with a clarity score greater than 6 from the preprocessed image data to obtain training image data.

3. The method for improving the facial quality of a person in an image according to claim 1, characterized in that: The step 2 is specifically as follows: all images in the training image data are processed using the cv2.resize function and the INTER_AREA interpolation method, and each image is processed to a resolution of 512×512; the images with a resolution of 512×512 are randomly horizontally flipped and color-jittered using the random_flip function and the random_color_jitter function, and then the pixel values ​​are normalized from [0,255] to [-1,1] using the normalize_pixels function to obtain the training image processing data.

4. The method for improving the facial quality of a person in an image according to claim 1, characterized in that: The step 3 is specifically as follows: construct a low-rank fine-tuning model: first, decompose the weight matrix W0 of UNet into the form of W0+BA, where B∈R^(d×r), A∈R^(r×d), r is set to 8, and d is the dimension of the input feature; design three independent fine-tuning branches, each of which contains an independent BA matrix pair, and the three independent fine-tuning branches are respectively used to focus on texture enhancement, light and shadow optimization, and expression naturalness; each branch receives the input image x, and passes W0+B i A i Feature extraction and transformation are performed in the form of, where i represents the branch index; Set batch_size=32, input the training image processing data into the initial low-rank fine-tuning model, perform model training, and obtain the trained low-rank fine-tuning model; The training adopts an alternating optimization strategy; Fix the W0 matrix and train the BA matrix; then adjust the weights of each branch according to the loss function evaluation results, and stop training when the loss function reaches the set threshold; Use gradient penalty and batch normalization techniques to stabilize the training process; All convolutional layers use 3×3 or 1×1 convolution kernels, the activation function uses LeakyReLU, α=0.2; set the interval time to check the gradient norm.

5. The method for improving the facial quality of a person in an image according to claim 1, characterized in that: The step 4 is specifically as follows: inputting the image with the face into the pre-trained FaceNet model, extracting the face features, and obtaining a 512-dimensional face feature vector; inputting the 512-dimensional face feature vector into the diffusion model, and fine-tuning the diffusion model through the trained low-rank fine-tuning model to generate an intermediate image with a resolution of 1024×1024, and the intermediate image is enlarged to a resolution of 2048×2048 through a 2x upsampling network output to obtain an enhanced facial feature image.

6. A device for improving the facial quality of a person in a picture, characterized by: include: The data screening module selects the required images from the face library as training image data; A data processing module processes the training image data to obtain an image of a set resolution as training image processing data; Construct a training module, construct a low-rank fine-tuning model, input the training image processing data into the initial low-rank fine-tuning model, perform model training, and obtain a trained low-rank fine-tuning model; The feature enhancement module extracts facial features from images with faces and obtains facial feature vectors; the facial feature vectors are input into the diffusion model, and the diffusion model is fine-tuned through the trained low-rank fine-tuning model to generate an enhanced facial feature image.

7. The device for improving the facial quality of a person in a picture according to claim 6, characterized in that: The data screening module is specifically as follows: first, the detection function is executed using the yolov5face algorithm to identify and screen out the face part of the front view; the screened face image is used to use the calculate_face_area function to ensure that the face area occupies 60%-80% of the center area of ​​the image; the calculate_rotation function is used to check whether the side rotation angle is within ±5°; then, the color histogram analysis is performed through the analyze_histogram function to eliminate abnormal exposure samples, ensure that the standard deviation of the three RGB channels of the image is less than 45, and the brightness is concentrated in the range of 0.4-0.8, and obtain the pre-processed image data; The preprocessed image data is screened using the select_high_quality_images function, which first uses cv2.GaussianBlur to perform Gaussian blur denoising on the image; then uses cv2.cvtColor to convert the image into a grayscale image, and then uses cv2.Laplacian(gray,cv2.CV_64F) to calculate the Laplacian operator, takes the absolute value of the result, and converts it to uint8 type, and finally calculates the average of all results to obtain a clarity score, sets the threshold to 6, and screens out 10,000 images with a clarity score greater than 6 from the preprocessed image data to obtain training image data.

8. The device for improving the facial quality of a person in a picture according to claim 6, characterized in that: The data processing module is specifically as follows: all images in the training image data are processed using the cv2.resize function and the INTER_AREA interpolation method, and each image is processed to a resolution of 512×512; the images with a resolution of 512×512 are randomly horizontally flipped and color-jittered using the random_flip function and the random_color_jitter function, and then the pixel values ​​are normalized from [0,255] to [-1,1] using the normalize_pixels function to obtain the training image processing data.

9. The device for improving the facial quality of a person in a picture according to claim 6, characterized in that: The construction of the training module is specifically as follows: construct a low-rank fine-tuning model: first, decompose the weight matrix W0 of UNet into W0+BA form, where B∈R^(d×r), A∈R^(r×d), r is set to 8, and d is the dimension of the input feature; design three independent fine-tuning branches, each of which contains an independent BA matrix pair, and the three independent fine-tuning branches are used to focus on texture enhancement, light and shadow optimization, and expression naturalness improvement; each branch receives the input image x, and passes W0+B i A i Feature extraction and transformation are performed in the form of, where i represents the branch index; Set batch_size=32, input the training image processing data into the initial low-rank fine-tuning model, perform model training, and obtain the trained low-rank fine-tuning model; The training adopts an alternating optimization strategy; Fix the W0 matrix and train the BA matrix; then adjust the weights of each branch according to the loss function evaluation results, and stop training when the loss function reaches the set threshold; Use gradient penalty and batch normalization techniques to stabilize the training process; All convolutional layers use 3×3 or 1×1 convolution kernels, the activation function uses LeakyReLU, α=0.2; set the interval time to check the gradient norm.

10. The device for improving the facial quality of a person in a picture according to claim 6, characterized in that: The feature enhancement module specifically includes: inputting an image with a face into a pre-trained FaceNet model to extract facial features and obtain a 512-dimensional facial feature vector; inputting the 512-dimensional facial feature vector into a diffusion model, and fine-tuning the diffusion model through a trained low-rank fine-tuning model to generate an intermediate image with a resolution of 1024×1024, and the intermediate image is enlarged to a resolution of 2048×2048 through a 2x upsampling network output to obtain an enhanced facial feature image.