Image transfer drawing method based on face recognition

By integrating face recognition algorithms and large-model fine-tuning technology in image processing, the automatic transformation and painting of face images is achieved, solving the shortcomings of automation and high efficiency in the existing technology, and improving the effect and efficiency of image processing.

CN120070161APending Publication Date: 2025-05-30CHENGDU TENGMU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149188.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has insufficient automation and high efficiency in facial recognition and image transfer, and it is difficult to achieve automated and efficient image processing.

Method used

By calling face recognition algorithms to detect faces in pictures, and combining Checkpoints big model, Lora fine-tuning model and Controlnet control network to convert face images to achieve automated image translation.

Benefits of technology

It realizes high-precision detection and stylized transformation of face images, improves user experience and image processing efficiency, and is suitable for entertainment and artistic creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070161A_ABST
    Figure CN120070161A_ABST
Patent Text Reader

Abstract

The invention relates to an image transfer method based on face recognition, and belongs to the field of image processing, and the method comprises the steps: 1, a user uploads a picture; 2, calling a face recognition algorithm to detect the picture; step 3, when a human face is detected, performing base64 coding on the picture; step 4, performing face focus cutting on the picture according to a proportion required for converting the theme, and outputting a cut face image for storage for later use; step 5, carrying out transduction processing on the face image; and step 6, storing the output picture and returning the picture to the user. The method aims at converting the recognized face image into the conversion image with the specific style or expression in real time, the user experience and the image processing efficiency are improved, high-precision face detection, feature extraction and stylized conversion are achieved, and the method is suitable for multiple fields such as entertainment and artistic creation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image transfer method based on face recognition. Background Art

[0002] Traditional face recognition technology mainly focuses on face authentication and ignores the creative processing of face images. At the same time, existing image transfer methods often require manual parameter adjustment, making it difficult to achieve automated and efficient transfer processing. Summary of the Invention

[0003] To solve the above problems of the prior art, the present invention provides an image transfer method based on face recognition.

[0004] An image transfer method based on face recognition includes the following steps:

[0005] Step 1: The user uploads a picture;

[0006] Step 2: Invoke the face recognition algorithm to detect the picture;

[0007] Step 3: When a face is detected, sort the picture, add a sequential annotation to the picture with a face, and perform base64 encoding on the picture;

[0008] Step 4: According to the ratio required by the transfer theme, divide the boundary of the picture pytorch tensor according to the obtained cropping size, complete the face focus cropping of the picture, and output the cropped face image;

[0009] Step 5: Perform transfer processing on the face image through the Checkpoints large model, the Lora fine-tuning model, and the Controlnet control network to obtain a transfer effect picture of a specified style:

[0010] Step 6: Save the output transfer effect picture and return it to the user.

[0011] Further, the specific steps of Step 1 include:

[0012] Step 101: The user uploads a picture;

[0013] Step 102: Perform server file transfer and storage according to the picture uploaded by the user;

[0014] Step 103: Obtain the file path and load the image node, and convert the picture into a pytorch tensor;

[0015] Step 104: Pass it to the subsequent node in the form of data of the pytorch tensor dimension information.

[0016] Further, the PyTorch tensor is represented as torch.size([Batch, Height, Width, Channel]), where Batch represents the batch size, Height represents the image height, Width represents the image width, and Channel represents the number of channels.

[0017] Further, the sorting method in step 3 includes:

[0018] Sorting in ascending or descending order based on the x-axis of the face bounding box of the image;

[0019] Sorting in ascending or descending order based on the y-axis of the face bounding box of the image;

[0020] Sorting in ascending or descending order based on the area of the face bounding box of the image.

[0021] Further, the specific process of cropping in step 4 includes:

[0022] Obtaining image cropping information according to the ratio required by the transfer drawing theme, including the short side cropping ratio, the long side cropping ratio, and the direction ratio threshold;

[0023] Storing the image information in the form of a high-dimensional array, where the high-dimensional array contains the image batch, the image width, the image height, and the RGB values at each position;

[0024] Obtaining the height and width of the original image to get the aspect ratio of the original image size, comparing the aspect ratio with the direction ratio threshold. If it is less than or equal to the threshold, the original image is determined to be a vertical image; otherwise, it is determined to be a horizontal image, thereby determining the cropping direction;

[0025] Obtaining the reserved sides, specifically: pre-calculating the width and height, and comparing the pre-calculated width and height with the width and height of the original image respectively, thereby obtaining the reserved sides, that is, obtaining the width and height after cropping;

[0026] Obtaining the center position of the cropping focus, specifically: obtaining the high-dimensional array data of the image information, processing the face coordinates according to the high-dimensional array data, and taking the part with the largest face area as the core face to calculate the center point coordinates;

[0027] Calculating the upper left and lower right coordinates of the cropping area according to the center position of the cropping focus and the cropped width and height, thereby generating a corresponding cropping box and cropping the image.

[0028] Further, the specific process of pre-calculating the width and height is: the pre-calculated width is the height multiplied by the short side cropping value and then divided by the long side cropping value; the pre-calculated height is the original image width multiplied by the long side cropping value and then divided by the short side cropping value.

[0029] Furthermore, step 5 specifically includes the following sub-steps:

[0030] Step 501: Reverse-infer the image prompt words;

[0031] Step 502: Adjust the parameters of the face image through the Checkpoints large model to generate images with different styles and qualities, and fine-tune the sampling of the Checkpoints large model through the Lora fine-tuning model;

[0032] Step 503: Constrain the sampling range through the positive and negative prompt words to ensure that the sampling map is sampled within the preset range;

[0033] Step 504: Combine the result of the reverse inference of the image prompt words with the range condition constraint to obtain the sampling constraint result;

[0034] Step 505: Use the Controlnet control network to preprocess the model algorithm for the pose and depth, and constrain the control of the subsequent large model sampling;

[0035] Step 506: Control the style adaptation according to the image prompt words to ensure that the style of the output image conforms to the transfer painting theme;

[0036] Step 507: Constrain the sampling through the KSampler sampler according to the image prompt words, sampling constraint result, Lora fine-tuning model, style control, and Controlnet control network content;

[0037] Step 508: Perform secondary sampling through the KSampler sampler to make the sampling details more abundant;

[0038] Step 509: Perform VAE decoding on the sampling result to decode the sampling noise and obtain the transfer painting effect diagram.

[0039] The beneficial effects of the present invention are reflected in that the present invention aims to convert the recognized face image into a transfer painting image with a specific style or expression in real time, improving the user experience and image processing efficiency; by integrating advanced face recognition algorithms and image processing technologies, high-precision face detection, feature extraction, and stylized transfer painting are realized, which are applicable to multiple fields such as entertainment and artistic creation. Brief Description of the Drawings

[0040] Figure 1 It is a flowchart of the method provided by the present invention. Detailed Embodiment

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] Embodiment 1: Refer to Figure 1 As shown, a method for image transfer and rendering based on face recognition in an embodiment of the present invention includes the following steps:

[0043] Step 1: The user uploads a picture;

[0044] Step 2: Invoke the face recognition algorithm to detect the picture;

[0045] Step 3: When a face is detected, sort the picture, add a sequential annotation to the picture with a face, and perform base64 encoding on the picture; Base64 encoding is used to unify the underlying format of the picture during transmission, facilitating subsequent operations such as logically storing files, logically converting torch picture tensors, etc.;

[0046] Step 4: According to the required ratio of the transfer and rendering theme, divide the boundaries of the picture pytorch tensor according to the obtained cropping size, complete the face focus cropping of the picture, and output the cropped face image;

[0047] Step 5: Perform transfer and rendering processing on the face image through the Checkpoints large model, the Lora fine-tuning model, and the Controlnet control network to obtain a transfer and rendering effect picture in a specified style:

[0048] Step 6: Save the output transfer and rendering effect picture and return it to the user.

[0049] Furthermore, the specific content of Step 1 includes:

[0050] Step 101: The user uploads a picture;

[0051] Step 102: Perform server file transfer and storage according to the picture uploaded by the user;

[0052] Step 103: Obtain the file path and the image loading node, and convert the picture into a pytorch tensor;

[0053] Step 104: Transmit it to the subsequent node in the form of the data of the pytorch tensor dimension information.

[0054] Among them, the PyTorch tensor is represented as torch.size([Batch, Height, Width, Channel]), where Batch represents the batch size, Height represents the image height, Width represents the image width, and Channel represents the number of channels. The image is transmitted as a PyTorch tensor to each node for processing. The input parameter Image of the workflow is uniformly PyTorch tensor data. During the face recognition process, the model parameters need to match the PyTorch image tensor data as the standard. After analysis by the internal neural network of the model, face information data is finally output;

[0055] Further, the sorting method in step 3 includes:

[0056] Sorting in ascending or descending order according to the x-axis of the face bounding box of the image;

[0057] Sorting in ascending or descending order according to the y-axis of the face bounding box of the image;

[0058] Sorting in ascending or descending order according to the area of the face bounding box of the image.

[0059] Among them, the specific process of cropping includes:

[0060] 1. Obtain image cropping information according to the ratio required by the transfer painting theme, including the short side ratio (float data), the long side ratio (float data), and the direction ratio threshold (float data) of the cropping; the short side ratio is the proportion of the short side of the cropped image, and the long side ratio is the proportion of the long side of the cropped image; this cropping does not limit width and height. The short side ratio can be width or height. The main judgment basis is the direction ratio threshold parameter, which affects whether the final cropping result of an image is a vertical image or a horizontal image; among them, the ratio required by the transfer painting theme can be obtained according to the ratio of the transfer painting theme ratio to the face cropping ratio of 2:3 to obtain the short side ratio and the long side ratio of the image cropping. By constraining the cropping ratio to be consistent with it, it is ensured that the theme display effect remains consistent in proportion when transitioning from the cropped image to the output size image of the transfer painting theme, and the transition effect is seamless.

[0061] 2. Store the image information in the form of a high-dimensional array, and the high-dimensional array includes the image batch, the image width, the image height, and the RGB values at each position;

[0062] 3. Obtain the height and width of the original image, obtain the aspect ratio of the original image size, compare the aspect ratio with the direction ratio threshold. If it is less than or equal to the threshold, it is determined that the original image is a vertical image, otherwise it is determined to be a horizontal image, so as to determine the cropping direction;

[0063] 4. Obtain the retained edges. Specifically: Pre-calculate the width and height, and compare the pre-calculated width and height with the width and height of the original image respectively. If the pre-calculated width is less than or equal to the width of the original image, the new cropped width = the pre-calculated width, and the new cropped height = the height of the original image; when the pre-calculated width is greater than the width of the original image, that is, retain the width of the original image and crop the height. In the case of cropping into a horizontal image, pre-calculate the height and then perform a pre-height comparison. calc_height = width * crop_rate_short / / crop_rate_long. If the pre-calculated height is less than or equal to the height of the original image, the new cropped height = the pre-calculated height, and the new cropped width = the width of the original image. When the calculated height is greater than the height of the original image, that is, retain the height of the original image, so as to obtain the retained edges, that is, obtain the cropped width and height. Among them, the specific process of pre-calculating the width and height is: The pre-calculated width is the height multiplied by the cropped short side value and then divided by the cropped long side value; the pre-calculated height is the width of the original image multiplied by the cropped long side value and then divided by the cropped short side value;

[0064] 5. Obtain the center position of the cropping focus. Specifically: Obtain the center position of the cropping focus. Specifically: Obtain the high-dimensional array data of the image information, process the face coordinates according to the high-dimensional array data, and take the part with the largest face area as the core face calculation center point coordinates; that is, the starting point start_x coordinate + width / 2 is equal to the face center point x coordinate, and the starting point start_y coordinate + height / 2 is equal to the face center point y coordinate;

[0065] In special cases (when no face is detected in the original image), the center point is the center of the original image. When no face is detected in the source image, the cropping algorithm performs proportional cropping with the coordinates of the center point (width / 2, height / 2) of the original image. The purpose of doing this is to still return the cropped image with the specified ratio to the client in the case of no detected face. Such an effect will solve the situation that the image without a face can still be proportionally cropped and returned to the next-level service for processing without throwing an exception or stopping due to subsequent inability to process;

[0066] 6. Calculate the upper left and lower right coordinates of the cropping area according to the center position of the cropping focus and the cropped width and height, so as to generate the corresponding cropping frame and crop the image.

[0067] In this embodiment, step 5 specifically includes the following sub-steps:

[0068] Step 501: Reverse infer the image prompt; The Vision + LLM vision-language model can be used to reverse infer the image prompt to obtain the prompt output;

[0069] Step 502: Adjust the parameters of the face image through the Checkpoints large model to generate images with different styles and qualities, and fine-tune the sampling of the Checkpoints large model through the Lora fine-tuning model;

[0070] Step 503: Constrain the sampling range through positive and negative prompts to ensure that the sampled images are sampled within the preset range;

[0071] Step 504: Combine the result of the reverse inference of the image prompt with the range condition constraint to obtain the sampling constraint result;

[0072] Step 505: Use the Controlnet control network object to preprocess the model algorithm for pose and depth, and constrain the control of subsequent large model sampling; for example: use the Zoe depth preprocessor to process the cropped image and load the depth model, or use the Canny preprocessor to process the cropped image and load the Canny model to complete the concatenation of the previous Positive&NegativeCondition;

[0073] Step 506: Control the style adaptation according to the image prompt to ensure that the style of the output image conforms to the transfer painting theme; for example: use IPAdapter to control the style algorithm;

[0074] Step 507: Use the KSampler sampler to perform constrained sampling according to the image prompt, sampling constraint result, Lora fine-tuning model, style control, and Controlnet control network content;

[0075] Step 508: Perform secondary sampling through the KSampler sampler to make the sampling details more abundant;

[0076] Step 509: Perform VAE decoding on the sampling result to decode the sampling noise and obtain the transfer painting effect picture.

[0077] In this embodiment, KSampler is a core sampler in ComfyUI. It allows users to generate images with different styles and qualities by adjusting various parameters. It is responsible for extracting samples from the latent space and converting them into high-quality images. This process involves reconstructing the image details to ensure that the finally generated images are clear and rich in details; at the same time, KSampler also supports a variety of sampling algorithms and schedulers, allowing users to select different generation strategies according to their needs.

[0078] In this embodiment, for the calculation of the upper left corner and lower right corner coordinates of the cropping area: ensure that the coordinates are integers and greater than 0. In an image with a face, use the center point of the face, that is, the coordinates of Face(FaceAreaWidth / 2,FaceAreaHeight / 2) as the cropping center point. If there is no face in the image, use (source image width / 2, source image height / 2) as the cropping center point. If the original image width is equal to the new cropping width, then (x1,y1)=(0,max(int(center_y–new_height / / 2),0)), (x2,y2)=(new_width,y1+new_height). If y2 is greater than the original image height (i.e., the cropping position has exceeded the original image area), then y1=height–new_height, y2=height; if the original image width is not equal to the new cropping height, that is, keep the height and crop the width, (x1,y1)=(max(int(center_x-new_width / / 2),0)), (x2,y2)=(x1+new_width,new_height). If x2 is greater than the original image width (i.e., the cropping exceeds the original image area), then x1=width-new_width, x2=width; finally, the cropped image cropped_image = image[:,y1:y2,x1:x2,:].

[0079] In the description of the embodiments of the present invention, the terms "first", "second", "third", "fourth" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third", "fourth" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more than two.

[0080] In the description of the embodiments of the present invention, specific features, structures, materials or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0081] In the description of the embodiments of the present invention, it should be understood that "-" and "~" represent the range between two numerical values, and this range includes the endpoints. For example: "A-B" represents the range greater than or equal to A and less than or equal to B. "A~B" represents the range greater than or equal to A and less than or equal to B.

[0082] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image transfer method based on face recognition, characterized in that: The following steps are involved: Step 1: User uploads a picture; Step 2: Call the face recognition algorithm to detect the image; Step 3: When a face is detected, the image is sorted, a sequence annotation is added to the image with the face, and the image is base64 encoded; Step 4: Calculate the cropping information according to the required ratio of the transferred subject, divide the boundary of the image pytorch tensor according to the obtained cropping size, complete the face focus cropping of the image, and output the cropped face image; Step 5: Use the Checkpoints large model, Lora fine-tuning model, and Controlnet control network to redraw the face image and obtain the redrawing effect of the specified style: Step 6: Save the output transfer rendering and return it to the user.

2. The image transfer method based on face recognition according to claim 1, characterized in that: The step 1 specifically includes: Step 101: User uploads a picture; Step 102: Transferring the server file according to the picture uploaded by the user; Step 103: Get the file path and load the image node to convert the image into a pytorch tensor; Step 104: Pass the data in the form of pytorch tensor dimension information to subsequent nodes.

3. The image transfer method based on face recognition according to claim 2, characterized in that: The pytorch tensor is represented as torch.size([Batch,Height,Width,Channel]), where Batch represents the batch size, Height represents the image height, Width represents the image width, and Channel represents the number of channels.

4. The image transfer method based on face recognition according to claim 1, characterized in that: The method for sorting in step 3 includes: Sort the face bounding boxes of the images in ascending or descending order based on their x-axis. Sort the face bounding boxes of the images in ascending or descending order based on their y-axis. Sort the images by the area of ​​the face bounding boxes from small to large or from large to small.

5. The image transfer method based on face recognition according to claim 1, characterized in that: The specific process of cutting in step 4 includes: Obtain image cropping information according to the required ratio of the transferred subject, including cropping short side ratio, cropping long side ratio and direction ratio threshold; Store the image information in the form of a high-dimensional array, which contains image batches, image width, image height, and RGB values ​​at each position; Get the height and width of the original image, get the aspect ratio of the original image size, compare the aspect ratio with the direction ratio threshold, if it is less than or equal to the threshold, the original image is judged to be a vertical image, otherwise it is judged to be a horizontal image, so as to determine the cropping direction; Obtain the retained edge, specifically: pre-calculate the width and height, and compare the pre-calculated width and height with the width and height of the original image respectively, so as to obtain the retained edge, that is, obtain the width and height after cropping; Obtaining the center position of the cropping focus, specifically: obtaining high-dimensional array data of the image information, processing the face coordinates according to the high-dimensional array data, and taking the largest part of the face area as the core face to calculate the center point coordinates; The coordinates of the upper left corner and lower right corner of the cropping area are calculated according to the center position of the cropping focus and the width and height of the cropping, so as to generate the corresponding cropping frame and crop the image.

6. The image transfer method based on face recognition according to claim 5, characterized in that: The specific process of precalculating the width and height is as follows: the precalculated width is the height multiplied by the cropping short side value, and then divided by the cropping long side value; the precalculated height is the original image width multiplied by the cropping long side value, and then divided by the cropping short side value.

7. The image transfer method based on face recognition according to claim 1, characterized in that: The step 5 specifically includes the following sub-steps: Step 501: reverse inferring the image prompt words; Step 502: Adjust the parameters of the face image through the Checkpoints large model to generate images of different styles and qualities, and fine-tune the sampling of the Checkpoints large model through the Lora fine-tuning model; Step 503: Using the positive prompt word and the reverse prompt word to constrain the sampling range, so as to ensure that the sampling graph is sampled within the preset range; Step 504: combining the image prompt word inverse deduction result with the range condition constraint to obtain a sampling constraint result; Step 505: Use the posture of the Controlnet control network object to perform in-depth model algorithm preprocessing to constrain the control of subsequent large model sampling; Step 506: performing style adaptation control according to the image prompt words to ensure that the style of the output image is consistent with the transfer theme; Step 507: Perform constrained sampling according to the image prompt words, sampling constraint results, Lora fine-tuning model, style control and Controlnet control network content through the KSampler sampler; Step 508: Perform secondary sampling through the KSampler sampler to make the sampling details richer; Step 509: Perform VAE decoding on the sampling result, decode the sampling noise, and obtain a redrawing effect diagram.