Head portrait generation method, system and device and storage medium
By receiving and pre-detecting multiple photos of users, training portrait models, and combining generative large models and style models to generate doctors' avatars, solving the cumbersome problems of doctors taking avatar photos, realizing personalized customization and efficient avatar generation.
Patent Information
- Application Number
- CN202311844109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the medical plastic surgery industry, doctors have difficulty taking photos of their avatars that meet the requirements due to busy work, resulting in a lack of personalized customized avatar solutions.
A method of avatar generation is proposed. By receiving multiple photos of users, pre-detection and portrait small model training, combining a generative large model and a style small model, the user's avatar template picture is generated based on text description, and finally the target avatar picture is obtained.
It realizes that while ensuring the custom avatar effect, the flexibility of AIGC is fully utilized, the workload is reduced, and the accuracy of image generation and user experience are improved.
Smart Images

Figure CN120235986A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a method, system, electronic device, and computer-readable storage medium for generating avatars. Background Art
[0002] In recent years, with the continuous improvement of people's living standards and the increasing demand for medical insurance among the general public, the medical plastic surgery and beauty industry has also been continuously developing, and modern medical aesthetic technologies are becoming increasingly advanced.
[0003] In the medical plastic surgery and beauty industry, doctor headshot photos in specific backgrounds and forms are often required for scenarios such as presenting doctor information to customers. However, most doctors are very busy with their work on a daily basis and may not be able to specifically take photos that meet the requirements. Therefore, an effective avatar generation solution is needed for personalized customization of doctor avatars. Summary of the Invention
[0004] The main purpose of this application is to propose a method, system, electronic device, and computer-readable storage medium for generating avatars, aiming to solve the problem of how to achieve personalized customization of user avatars.
[0005] In a first aspect, an embodiment of this application provides a method for generating an avatar, the method including:
[0006] Receiving multiple photos of the same user;
[0007] Performing pre-detection on the multiple photos of the user to determine whether the avatar information is complete;
[0008] In the case where the avatar information in the multiple photos is incomplete, training a portrait small model corresponding to the user avatar based on the multiple photos;
[0009] Selecting a portrait style small model according to the required scenario;
[0010] Obtaining a text description of the avatar picture to be generated, and combining a generative large model, the portrait small model, and the portrait style small model to generate a corresponding avatar template picture of the user according to the text description;
[0011] Obtaining a target avatar picture based on the avatar template picture.
[0012] In a second aspect, an embodiment of this application provides an avatar generation system, the system including:
[0013] A receiving module, configured to receive multiple photos of the same user;
[0014] A detection module, configured to perform pre-detection on the multiple photos of the user to determine whether the avatar information is complete;
[0015] A training module, configured to train a portrait mini-model corresponding to the user portrait according to the multiple photos in a case where the portrait information in the multiple photos is incomplete;
[0016] A selection module, configured to select a portrait style mini-model according to a required scenario;
[0017] A generation module, configured to obtain a text description of a portrait picture to be generated, combine a generative large model, the portrait mini-model, and the portrait style mini-model, generate a portrait template picture corresponding to the user according to the text description, and obtain a target portrait picture according to the portrait template picture.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes: a memory, a processor, and a portrait generation program stored on the memory and executable on the processor. When the portrait generation program is executed by the processor, the portrait generation method as described above is implemented.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where a portrait generation program is stored on the computer-readable storage medium. When the portrait generation program is executed by a processor, the portrait generation method as described above is implemented.
[0020] The portrait generation method, system, electronic device, and computer-readable storage medium provided by the embodiments of the present application can customize user portraits by scenario, and can give full play to the flexibility of AIGC on the premise of ensuring the effect of the customized portraits. Moreover, the embodiments of the present application combine and use three types of models, namely, a generative large model, a portrait mini-model, and a portrait style mini-model, and train a portrait mini-model each time multiple photos of the same user are received, which can also achieve the corresponding picture generation effect while reducing the workload. Description of the Drawings
[0021] The drawings here are used to provide a further understanding of the present application and constitute a part of the present application. It should be understood that these drawings only depict some embodiments disclosed according to the present application and should not be regarded as limiting the scope of the present application.
[0022] Figure 1 An application environment architecture diagram for implementing various embodiments of the present application;
[0023] Figure 2 A flowchart of a portrait generation method proposed in the first embodiment of the present application;
[0024] Figure 3 For Figure 2 A detailed flowchart diagram of step S202 in
[0025] Figure 4 Schematic diagram of an SDXL model in this application;
[0026] Figure 5 For Figure 2 Schematic diagram of a refined process for step S208 in;
[0027] Figure 6 Flowchart of an avatar generation method proposed in the second embodiment of this application;
[0028] Figure 7 Another form of flowchart of the avatar generation method proposed in the second embodiment of this application;
[0029] Figure 8 Schematic diagram of the hardware architecture of an electronic device proposed in the third embodiment of this application;
[0030] Figure 9 Schematic diagram of the modules of an avatar generation system proposed in the fourth embodiment of this application
[0031] Figure 10 Schematic diagram of the modules of an avatar generation system proposed in the fifth embodiment of this application. Detailed implementation manners
[0032] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope protected by this application.
[0033] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of this application are only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0034] The following provides explanations of the terms involved in this application:
[0035] Artificial Intelligence Generated Content (AIGC): It refers to the technical methods of artificial intelligence such as generative adversarial networks and large pre-trained models. Through the learning and recognition of existing data, it is a technology that can generate relevant content with appropriate generalization ability. The core idea of AIGC technology is to use artificial intelligence algorithms to generate content with a certain degree of creativity and quality. Through training models and learning a large amount of data, AIGC can generate content related to the input conditions or instructions. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, etc. Along with the development of AIGC technology, customized user images have also become more mature.
[0036] Stable Diffusion (SD) model: A deep learning text-to-image generation model, which is a type of AIGC model and is mainly used to draw imaginary pictures based on text descriptions. It adopts a more stable and controllable diffusion process, thus being able to generate high-quality images. By inputting text prompts, the model will output an image that matches the prompts.
[0037] Prompt: It is the input of the AIGC model used to draw images. Generally, Prompts are divided into positive prompts and negative prompts.
[0038] Image-to-Image: It is a way of image generation by the AIGC model. Different from only using text prompts, the Image-to-Image technology uses a reference image and text prompts as common inputs to perform secondary creation on the original image.
[0039] SDXL model: The latest optimized version of the SD model, which is a two-stage cascaded diffusion model, including a Base model and a Refiner model. The main work of the Base model is the same as that of the SD model, and it has capabilities such as text-to-image, image-to-image, and inpainting. After the Base model, the Refiner model is cascaded to refine the latent features of the images generated by the Base model. Essentially, it is doing the work of image-to-image.
[0040] LoRA model (Low-Rank Adaptation of Large Language Models): It can be understood as a plugin for the SD model. Without modifying the SD model, it is a small model formed by training a small number of pictures. It can be used in combination with the large model to interfere with the results generated by the large model to achieve customized requirements, and the training resources required are much smaller than training the SD model.
[0041] VAE (Variable Auto Encoder): A generative model that can provide the key ability to efficiently extract the latent features of data.
[0042] The technical solutions of the present application will be specifically described below in conjunction with each embodiment.
[0043] Please refer to Figure 1 , Figure 1 FIG. is an application environment architecture diagram for implementing each embodiment of the present application. The present application can be applied to an application environment including, but not limited to, a client 2, a server 4, and a network 6.
[0044] Among them, the client 2 is used to receive multiple photos uploaded by the user, receive text descriptions, and display the finally generated avatar pictures to the user, etc. The server 4 is used to train a portrait LoRA small model based on the multiple photos, and based on the text description, the generative large model, the LoRA portrait small model, and the selected portrait style LoRA small model, generate an avatar template picture corresponding to the user, and then obtain the final avatar picture by performing face fusion on the avatar template picture and the multiple photos respectively.
[0045] The client 2 can be a terminal device such as a PC (Personal Computer), a mobile phone, a tablet computer, a portable computer, a wearable device, etc. The server 4 can be a computing device such as a rack-mounted server, a blade server, a tower server, or a cabinet server, and can be an independent server or a server cluster composed of multiple servers.
[0046] The network 6 can be an enterprise internal network (Intranet), the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi and other wireless or wired networks. The server 4 and one or more of the clients 2 are communicatively connected through the network 6 for data transmission and interaction.
[0047] Embodiment 1
[0048] Such as Figure 2As shown in the figure, it is a flowchart of a method for generating an avatar proposed in the first embodiment of the present application. It can be understood that the flowchart in the embodiment of this method is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. Hereinafter, the server is used as the execution subject to illustrate this method.
[0049] This method includes the following steps:
[0050] S200, Receive multiple photos of the same user.
[0051] In this embodiment, according to multiple photos uploaded by the user, a required customized avatar image can be automatically generated through a model. The photos can be life photos taken by the user himself, such as photos taken by mobile phone selfies, etc., as long as it is ensured that the photos contain the face image of the same user. Generally, the number of the multiple photos can be 5 - 20.
[0052] S201, Perform pre - detection on the multiple photos of the user to determine whether the avatar information is complete.
[0053] In this embodiment, the pre - detection includes resolution and clarity detection, human body posture detection, and face detection. Perform resolution and clarity detection, human body posture detection, and face detection on the multiple photos respectively. When the resolution and clarity of at least one photo are both high, and the human body posture is complete, and the face part is also complete, it is determined that the avatar information of the multiple photos is complete. Among them, both the resolution and clarity being high means that the resolution and clarity of the photo respectively reach a preset threshold. The human body posture detection includes: identifying the human body posture in the photo, detecting whether the human body posture is a front posture, and if it is a front posture, it is determined that the human body posture of the photo is complete. The face detection includes: identifying the face part in the photo, detecting whether the face part reaches a preset proportion (such as 90%) of a complete face, that is, whether the photo is a frontal photo of the user. If it reaches the preset proportion, it is determined that the face part of the photo is complete.
[0054] In other alternative embodiments, it is also possible to first obtain the style requirements for the avatar image to be generated, such as clothing requirements, background requirements, human body posture requirements, etc. Then, on the basis of the above - mentioned pre - detection, further detect whether the multiple photos meet the style requirements. If they do not meet the style requirements, it still belongs to the situation where the avatar information is incomplete.
[0055] S202, In the case where the avatar information is incomplete, train a portrait LoRA small model corresponding to the user's avatar according to the multiple photos.
[0056] In this embodiment, each time a set of photos (multiple photos of the same user) is received, if the avatar information is incomplete, the portrait LoRA small model will be retrained based on the multiple photos. That is to say, the portrait LoRA small model is not fixed, but is separately learned and trained according to the photos of different users to obtain a customized small model exclusive to that user.
[0057] The generative large model is a diffusion model that will diverge according to the input text description to generate different pictures. The portrait LoRA small model is used to learn the appearance features of the user based on the multiple photos to limit the diffusion direction of the generative large model. That is to say, it makes the avatar pictures generated by the generative large model remain similar to the appearance of the user.
[0058] Specifically, further refer to Figure 3 , which is a refined flowchart of the above step S202. It can be understood that this flowchart is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. In this embodiment, the step S202 specifically includes:
[0059] S2020, preprocess the multiple photos to obtain a training set for training the portrait LoRA small model.
[0060] In this embodiment, the preprocessing includes but is not limited to image data augmentation, face attribute recognition, scene recognition, human pose recognition, expression recognition, and annotation of text labels. Specifically, the process of the preprocessing includes:
[0061] (1) Perform image data augmentation on the multiple photos. The image data augmentation refers to operations such as flipping, folding, and cropping the multiple photos. For example, changing a frontal face image to a side face image to expand the number and styles of photos and improve the generalization of the multiple photos.
[0062] (2) Perform face attribute recognition on the augmented photos. The face attribute recognition refers to identifying the key point information of the faces in the augmented photos, including the key points of the face, eyes, eyebrows, lips, and nose contour. Generally, the face attribute recognition of the photos can be achieved through 106 key point information. In addition, the face attribute recognition also includes various shooting perspectives and various degrees of freedom of the neck.
[0063] (3) Perform scene recognition, human pose recognition, and expression recognition on the augmented photos. The above various recognitions can be processed using any existing feasible recognition methods or models, which are not limited here.
[0064] (4) Based on the results of the face attribute recognition and the results of the scene, human body posture, and expression recognition, annotate text tags for the augmented photo. The annotated text tags refer to adding tags to the photo according to the face key point information in the augmented photo, as well as information such as the scene, human body posture, or character expression of the augmented photo, such as a dining scene, a smiling expression, etc.
[0065] S2022, train the LoRA model according to the training set to obtain the portrait LoRA small model.
[0066] First, set the model training parameters. In this embodiment, the model training parameters include picture resolution, picture size, the maximum number of training rounds, how many rounds to output once, etc. For the portrait LoRA small model, the maximum number of training rounds can be set to, for example, 20 times. Then, based on the training parameters, perform model training based on the training set.
[0067] The training set photos required for the portrait LoRA small model do not need to be too many, but the number of training rounds should be a little more. For multiple photos of the same user, the training process takes about 20 minutes. Therefore, even if the portrait LoRA small model needs to be retrained every time multiple photos of the same user are received, it will not take too long to wait.
[0068] Back to Figure 2 , S204, select the portrait style LoRA small model according to the required scene.
[0069] By pre-training different styles according to different scene requirements, multiple different portrait style LoRA small models can be obtained. When customizing the generation of avatar pictures, just select the corresponding portrait style LoRA small model according to the required scene, such as the style of a one-inch ID photo.
[0070] S206, obtain the text description of the avatar picture to be generated, and combine the generative large model, the portrait LoRA small model, and the portrait style LoRA small model to generate the avatar template picture corresponding to the user according to the text description.
[0071] The text description is used to describe the characteristics of the avatar picture to be generated, including size, background, clothing, expression, hairstyle, posture, perspective, etc. In an optional embodiment, the user can customize and describe the characteristics of the avatar picture they want, such as inputting "a professional woman wearing a red suit".
[0072] In another alternative embodiment, to lower the operation threshold, dropdown options for various avatar image features can also be provided on the front-end page, with multiple preset prompt words available for each feature, simplifying the user operation. The user only needs to select the corresponding target prompt word in the corresponding dropdown box according to the desired avatar image feature. Specifically, it includes:
[0073] (1) Provide dropdown options for various avatar image features on the front-end page, with multiple preset prompt words for each feature.
[0074] (2) Receive the target prompt words selected by the user in each option.
[0075] (3) Add the target prompt words to a preset prompt word template to generate a complete prompt word.
[0076] Since the user has selected multiple target prompt words for multiple avatar image features, they need to be integrated before inputting into the model, and a complete prompt word is generated through the prompt word template.
[0077] The generative large model can be the SD large model or other AIGC large models. In this embodiment, the SD large model is taken as an example for illustration. After obtaining the portrait LoRA small model and the portrait style LoRA small model, the SD large model combines the portrait LoRA small model and the portrait style LoRA small model to automatically generate the corresponding avatar template image for the user according to the text description.
[0078] In this embodiment, the SD large model uses the latest optimized version SDXL model. As Figure 4 shown, it is a schematic diagram of a SDXL model. In Figure 4 it, the SDXL model mainly includes parts such as the Base model, the Refiner model, and the VAE decoder (Decoder).
[0079] Among them, VAE is a generative model that can provide the key ability to efficiently extract the Latent features of data. When the input of the SDXL model is an image, first, the encoder structure of VAE is used to convert the input image into Latent features, and then the Base model and the Refiner model continuously optimize the Latent features. Finally, the decoder structure of VAE is used to reconstruct the Latent features into a pixel-level image. In addition to extracting Latent features and pixel-level reconstruction of images, VAE can also improve the high-frequency details, small object features, and overall image color in the generated images. When the input of the SDXL model is text, the encoder structure of VAE is not required, and only the decoder structure is needed for image reconstruction.
[0080] In this embodiment, the specific process of generating the avatar template image is as follows: First, input the custom text description / prompt, and generate Latent features through the VAE and Base models. Then, add a certain amount of noise (Noise) to the Latent features. On this basis, use the Refiner model to denoise to improve the overall quality and local details of the image. Finally, reconstruct the Latent features into a pixel-level image through the VAE decoder to obtain the avatar template image. During this process, it is also necessary to combine the portrait LoRA small model and the style LoRA small model to correct the generative large model (SDXL model). Then, generate the avatar template image corresponding to the user through the corrected model according to the text description.
[0081] Specifically, the process of the correction includes:
[0082] (1) Obtain the portrait feature parameters corresponding to the user based on the portrait LoRA small model.
[0083] Since the portrait LoRA small model is trained according to the multiple photos of the user and is a customized small model exclusive to the user, the portrait feature parameters of the portrait LoRA small model are parameters that conform to the appearance characteristics of the user.
[0084] (2) Obtain the required scene style parameters based on the portrait style LoRA small model.
[0085] The portrait style LoRA small model is selected according to the required scene, so the scene style parameters correspond to the scene to be generated.
[0086] (3) Adjust the weights of the corresponding parameters in the generative large model according to the portrait feature parameters and the scene style parameters.
[0087] Through the portrait feature parameters of the portrait LoRA small model and the scene style parameters of the portrait style LoRA small model, the weights of the corresponding parameters in the generative large model (SDXL model) can be increased, that is, the weights of the portrait feature parameters that conform to the appearance characteristics of the user and the scene style parameters that conform to the scene to be generated are increased, thereby restricting the diffusion direction of the generative large model (SDXL model) and customizing the portrait appearance and picture style.
[0088] Finally, through the corrected generative large model (SDXL model), the avatar template image corresponding to the user can be generated.
[0089] S208, obtain the target avatar image according to the avatar template image.
[0090] In this embodiment, the avatar template image can be directly used as the final target avatar image.
[0091] In a preferred embodiment, the target avatar image can also be obtained by fusing the avatar template image with the multiple photos.
[0092] Specifically, further refer to Figure 5 , which is a detailed flowchart of step S208 above. It can be understood that this flowchart is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted. In the preferred embodiment, step S208 specifically includes:
[0093] S2080, perform face fusion on the avatar template image and the multiple photos respectively to obtain multiple fused images.
[0094] In this embodiment, after the avatar template image is output by the SDXL model, it is also necessary to perform face fusion on the avatar template image and the multiple photos respectively to obtain multiple fused images. The fused images will contain the facial appearance features in the corresponding photos, as well as other appearance features and content in the avatar template image.
[0095] S2082, sort the multiple fused images according to the face similarity between the multiple fused images and the corresponding photos.
[0096] After the face fusion, the face similarity of each fused image and the corresponding photo will be compared, and the similarity result will be output. According to the face similarity between the fused image and the corresponding photo, the multiple fused images can be sorted accordingly.
[0097] S2084, obtain the fused image with the highest similarity according to the sorting result and output it as the final avatar image.
[0098] In this embodiment, the fused image with the highest face similarity to the corresponding photo will be used as the finally generated avatar image.
[0099] The finally generated avatar image can be uploaded to the cloud for storage, and the URL (Uniform Resource Locator) address of the cloud storage will be returned to the client.
[0100] The avatar generation method proposed in this embodiment can customize doctor avatars by scenario. On the premise of ensuring the effect of the customized avatar, it gives full play to the flexibility of AIGC. Moreover, this embodiment combines three types of models: a generative large model, a portrait LoRA small model, and a portrait style LoRA small model. By training a portrait LoRA small model each time multiple photos of the same user are received, the corresponding image generation effect can be achieved while reducing the workload. In addition, this embodiment adopts the method of fusing the avatar template image with the photos uploaded by the user to find the best avatar image, which can improve the accuracy of the output result and enhance the user experience.
[0101] Embodiment 2
[0102] As Figure 6 shown, it is a flowchart of an avatar generation method proposed in the second embodiment of this application. In the second embodiment, on the basis of the above first embodiment, the avatar generation method further includes step S410. It can be understood that the flowchart in the embodiment of this method is not used to limit the order of execution steps. According to needs, some steps in this flowchart can also be added or deleted.
[0103] The method includes the following steps:
[0104] S400, Receive multiple photos of the same user.
[0105] In this embodiment, according to multiple photos uploaded by the user, a customized avatar image required can be automatically generated by the model. The photos can be life photos taken by the user himself, such as photos taken by mobile phone selfies, etc., as long as it is ensured that the photos contain the face images of the same user. Generally, the number of the multiple photos can be 5 - 20.
[0106] S401, Perform pre - detection on the multiple photos of the user to determine whether the avatar information is complete. If the avatar information is incomplete, execute step S402; if the avatar information is complete, execute step S410.
[0107] In this embodiment, the pre - detection includes resolution and clarity detection, human body posture detection, and face detection. Perform resolution and clarity detection, human body posture detection, and face detection on the multiple photos respectively. When the resolution and clarity of at least one photo are both high, and the human body posture is complete and the face part is also complete, it is determined that the avatar information of the multiple photos is complete.
[0108] In other alternative embodiments, it is also possible to first obtain the style requirements for the avatar image to be generated, such as clothing requirements, background requirements, human pose requirements, etc. Then, on the basis of the above pre-detection, further detect whether the multiple photos meet the style requirements. If they do not meet the style requirements, it still belongs to the situation where the avatar information is incomplete.
[0109] S402. Train a portrait LoRA small model corresponding to the user avatar according to the multiple photos.
[0110] In this embodiment, each time a set of photos (multiple photos of the same user) is received, if the avatar information is incomplete, the portrait LoRA small model will be retrained according to the multiple photos. That is to say, the portrait LoRA small model is not fixed, but is separately learned and trained according to the photos of different users to obtain a customized small model exclusive to that user.
[0111] The portrait LoRA small model is used to learn the appearance features of the user according to the multiple photos to restrict the diffusion direction of the generative large model. That is to say, it makes the avatar image generated by the generative large model maintain similarity with the appearance of the user. For multiple photos of the same user, the training process takes about 20 minutes. Therefore, even if the portrait LoRA small model needs to be retrained each time multiple photos of the same user are received, it will not take too long to wait.
[0112] S404. Select a portrait style LoRA small model according to the required scenario.
[0113] By pre-learning and training different styles according to different scenario requirements, multiple different portrait style LoRA small models can be obtained. When customized avatar images need to be generated, just select the corresponding portrait style LoRA small model according to the required scenario, such as the style of a one-inch ID photo.
[0114] S406. Obtain a text description of the avatar image to be generated, and combine the generative large model with the portrait LoRA small model and the portrait style LoRA small model to generate a template avatar image corresponding to the user according to the text description.
[0115] The text description is used to describe the features of the avatar image to be generated, including size, background, clothing, expression, hairstyle, pose, perspective, etc. In an alternative embodiment, the user can customize and describe the features of the avatar image they want, such as inputting "a professional woman wearing a red suit".
[0116] In another alternative embodiment, in order to lower the operation threshold, dropdown options for various avatar picture features can also be provided on the front-end page, with multiple preset alternative prompt words provided for each feature to simplify user operations. The user only needs to select the corresponding target prompt word in the corresponding dropdown box according to the desired avatar picture feature.
[0117] The generative large model can be the SD large model or other AIGC large models. In this embodiment, the SD large model is taken as an example for illustration. After obtaining the portrait LoRA small model and the portrait style LoRA small model, the SD large model combines the portrait LoRA small model and the portrait style LoRA small model to automatically generate the avatar template picture corresponding to the user according to the text description.
[0118] In this embodiment, the SD large model adopts the latest optimized version SDXL model. Combining the portrait LoRA small model and the portrait style LoRA small model, the generative large model (SDXL model) is corrected. Then, the avatar template picture corresponding to the user is generated through the corrected model according to the text description. Specifically, the correction process includes:
[0119] (1) Obtain the portrait feature parameters corresponding to the user based on the portrait LoRA small model.
[0120] (2) Obtain the required scene style parameters based on the portrait style LoRA small model.
[0121] (3) Adjust the weights of the corresponding parameters in the generative large model according to the portrait feature parameters and the scene style parameters.
[0122] Through the portrait feature parameters of the portrait LoRA small model and the scene style parameters of the portrait style LoRA small model, the weights of the corresponding parameters in the generative large model (SDXL model) can be increased, that is, the weights of the portrait feature parameters that conform to the appearance characteristics of the user and the scene style parameters that conform to the required generated scene are increased, thereby restricting the diffusion direction of the generative large model (SDXL model) and customizing the portrait appearance and picture style.
[0123] Finally, through the corrected generative large model (SDXL model), the avatar template picture corresponding to the user can be generated.
[0124] S408. Obtain the target avatar picture according to the avatar template picture.
[0125] In this embodiment, the avatar template picture can be directly used as the final target avatar picture.
[0126] In a preferred embodiment, the target avatar image can also be obtained by fusing the avatar template image with the multiple photos. Specifically, first, the avatar template image is respectively face-fused with the multiple photos to obtain multiple fused images. Then, the multiple fused images are sorted according to the face similarity between the multiple fused images and the corresponding photos. Then, the fused image with the highest similarity is obtained according to the sorting result and output as the final avatar image.
[0127] S410, the final avatar image is obtained by adjusting and beautifying the multiple photos.
[0128] In this embodiment, the adjustment and beautification include image rotation and alignment, portrait skin beautification, Real-ESRGAN image enhancement operation, etc. Through the above adjustment and beautification, the original photos can meet the conditions of the required generated avatar image.
[0129] Among them, Real-ESRGAN is used for super-resolution enhancement of low-resolution images. Super-Resolution refers to the process of improving the resolution of the original image through hardware or software methods, and obtaining a high-resolution image from a series of low-resolution images. That is to say, the image is enlarged on the premise of keeping the clarity of the original image unchanged. Real-ESRGAN is a deep learning super-resolution model that realizes the super-resolution process of the original image through deep learning, and can also perform degradation processing on the data during data augmentation, and can also perform operations such as de-blurring, de-noising, and de-scratching during super-resolution.
[0130] If the user is not satisfied with the avatar image obtained through the above adjustment and beautification, they can also input a text description of the required generated avatar image and regenerate the target avatar image according to the process in the case where the avatar information is incomplete.
[0131] As Figure 7 shown, it is a flowchart of another form of the avatar generation method in this embodiment. Figure 7 The specific implementation processes of each step in
[0132] The avatar generation method proposed in this embodiment can customize doctor avatars by scenario, giving full play to the flexibility of AIGC while ensuring the effect of the customized avatar. Moreover, this embodiment combines three types of models: a generative large model, a portrait LoRA small model, and a style LoRA small model. By training a portrait LoRA small model each time multiple photos of the same user are received, the corresponding image generation effect can be achieved with reduced workload. In addition, this embodiment uses the method of fusing the avatar template image with the photos uploaded by the user to find the best avatar image, which can improve the accuracy of the output result and enhance the user experience. For the case where the avatar information in the multiple photos uploaded by the user is already relatively complete, this embodiment does not need to generate avatar images through the above models, but instead adds a beautification function for the original photos to ensure that the obtained avatar images better meet the requirements.
[0133] Embodiment III
[0134] As Figure 8 shown, FIG. shows a schematic hardware architecture of an electronic device 20 proposed in the third embodiment of the present application. In this embodiment, the electronic device 20 may include, but is not limited to, a memory 21, a processor 22, and a network interface 23 that are communicatively connected to each other through a system bus. It should be noted that Figure 8 only the electronic device 20 with components 21-23 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. In this embodiment, the electronic device 20 may be a server.
[0135] The memory 21 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 21 may be an internal storage unit of the electronic device 20, such as the hard disk or memory of the electronic device 20. In other embodiments, the memory 21 may also be an external storage device of the electronic device 20, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 20. Of course, the memory 21 may also include both the internal storage unit of the electronic device 20 and its external storage device. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed in the electronic device 20, such as the program code of the avatar generation system 60, etc. In addition, the memory 21 may also be used to temporarily store various data that have been output or will be output.
[0136] In some embodiments, the processor 22 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 22 is generally used to control the overall operation of the electronic device 20. In this embodiment, the processor 22 is used to run the program code stored in the memory 21 or process data, such as running the avatar generation system 60, etc.
[0137] The network interface 23 may include a wireless network interface or a wired network interface, and the network interface 23 is generally used to establish a communication connection between the electronic device 20 and other electronic devices.
[0138] Embodiment Four
[0139] As Figure 9 shown, a module schematic diagram of an avatar generation system 60 is proposed in the fourth embodiment of this application. The avatar generation system 60 can be divided into one or more program modules, and one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of this application. The program modules referred to in the embodiments of this application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment.
[0140] In this embodiment, the avatar generation system 60 includes:
[0141] A receiving module 600, configured to receive multiple photos of the same user.
[0142] A detection module 601, configured to pre-detect multiple photos of the user to determine whether the avatar information is complete.
[0143] A training module 602, configured to train a portrait LoRA small model corresponding to the user avatar according to the multiple photos in the case where the avatar information is incomplete.
[0144] A selection module 604, configured to select a portrait style LoRA small model according to the required scenario.
[0145] A generation module 606, configured to obtain a text description of the avatar picture to be generated, combine a generative large model, the portrait small model, and the portrait style small model, generate a template avatar picture corresponding to the user according to the text description, and obtain a target avatar picture according to the template avatar picture.
[0146] For the specific implementation process of the functions of each of the above modules, reference can be made to the description in the first embodiment above, and details are not described herein again.
[0147] Embodiment Five
[0148] As Figure 10 shown, this is a schematic diagram of the modules of an avatar generation system 60 proposed in the fifth embodiment of the present application. In this embodiment, in addition to including the receiving module 600, the detection module 601, the training module 602, the selection module 604, and the generation module 606 in the fifth embodiment, the avatar generation system 60 further includes an adjustment module 608.
[0149] The adjustment module 608 is configured to obtain a final avatar picture by adjusting and beautifying the multiple photos in the case where the avatar information in the multiple photos is complete.
[0150] For the specific implementation process of the function of the adjustment module 608, reference can be made to the description in the second embodiment above, and details are not described herein again.
[0151] Embodiment Six
[0152] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing an avatar generation program, and the avatar generation program can be executed by at least one processor to enable the at least one processor to execute the steps of the avatar generation method as described above.
[0153] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the avatar generation method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.
[0154] It should be noted that, in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.
[0155] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0156] Obviously, those skilled in the art should understand that the various modules or steps of the above embodiments of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0157] The above are only the preferred embodiments of the embodiments of the present application, and do not limit the patent scope of the embodiments of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the embodiments of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the embodiments of the present application.
Claims
1. A method for generating an avatar, characterized in that, The method includes: Receiving multiple photos of the same user; Pre-detecting the multiple photos of the user to determine whether the avatar information is complete; In the case where the avatar information in the multiple photos is incomplete, training a portrait small model corresponding to the user avatar based on the multiple photos; Selecting a portrait style small model according to the required scenario; Obtaining a text description of the avatar picture to be generated, and combining a generative large model, the portrait small model, and the portrait style small model to generate a corresponding avatar template picture of the user according to the text description; Obtaining a target avatar picture according to the avatar template picture.
2. The avatar generation method according to claim 1, wherein The obtaining a target avatar picture according to the avatar template picture includes: Performing face fusion on the avatar template picture and the multiple photos respectively to obtain multiple fusion pictures; Sorting the multiple fusion pictures according to the face similarity between the multiple fusion pictures and the corresponding photos; Obtaining the fusion picture with the highest similarity according to the sorting result and outputting it as the final avatar picture.
3. The avatar generation method according to claim 1 or 2, characterized in that, The combining a generative large model, the portrait small model, and the portrait style small model to generate a corresponding avatar template picture of the user according to the text description includes: Correcting the generative large model based on the portrait small model and the portrait style small model; Generating a corresponding avatar template picture of the user through the corrected model according to the text description.
4. The avatar generation method according to claim 3, wherein The correcting the generative large model based on the portrait small model and the portrait style small model includes: Obtaining portrait feature parameters corresponding to the user based on the portrait small model; Obtaining required scene style parameters based on the portrait style small model; Adjusting the weights of the corresponding parameters in the generative large model according to the portrait feature parameters and the scene style parameters.
5. The avatar generation method according to claim 1 or 2, characterized in that, The pre-detecting the multiple photos of the user to determine whether the avatar information is complete includes: Performing resolution and clarity detection, human pose detection, and face detection on the multiple photos respectively; Determining that the avatar information of the multiple photos is complete when the resolution and clarity of at least one photo reach a preset threshold, the required human pose is complete, and the face part is complete.
6. The avatar generation method according to claim 5, wherein The human pose detection includes: identifying the human pose in the photo, detecting whether the human pose is a front pose, and if it is a front pose, determining that the human pose of the photo is complete; The face detection includes: identifying the face part in the photo, detecting whether the face part reaches a preset ratio of a complete face, and if it reaches the preset ratio, determining that the face part of the photo is complete.
7. The avatar generation method according to claim 1 or 2, characterized in that, The training a portrait small model corresponding to the user avatar based on the multiple photos includes: Preprocessing the multiple photos to obtain a training set for training the portrait small model; Training a LoRA model according to the training set to obtain the portrait small model.
8. The avatar generation method according to claim 7, wherein The preprocessing the multiple photos includes: Performing image data augmentation on the multiple photos to expand the number and styles of the photos; Perform face attribute recognition on the augmented photo; Perform scene recognition, human pose recognition, and expression recognition on the augmented photo; According to the results of the face attribute recognition and the results of the scene, human pose, and expression recognition, label text tags for the augmented photo.
9. The avatar generation method according to claim 1 or 2, characterized in that, The method further includes: When the head portrait information of the multiple photos is complete, perform image rotation alignment, portrait beauty enhancement, and image enhancement operations on the multiple photos to obtain the final head portrait picture.
10. An avatar generation system, characterized in that, The system includes: A receiving module for receiving multiple photos of the same user; A detection module for pre-detecting the multiple photos of the user to determine whether the head portrait information is complete; A training module for training a portrait small model corresponding to the user's head portrait according to the multiple photos when the head portrait information of the multiple photos is incomplete; A selection module for selecting a portrait style small model according to the required scene; A generation module for obtaining a text description of the head portrait picture to be generated, combining a generative large model, the portrait small model, and the portrait style small model, generating a head portrait template picture corresponding to the user according to the text description, and obtaining a target head portrait picture according to the head portrait template picture.
11. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a head portrait generation program stored on the memory and executable on the processor. When the head portrait generation program is executed by the processor, it implements the head portrait generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, A head portrait generation program is stored on the computer-readable storage medium. When the head portrait generation program is executed by the processor, it implements the head portrait generation method according to any one of claims 1 to 9.