Image synthesis method, image synthesis model training method, medium and product
By constructing an image synthesis model combining portrait low-rank matrix and style low-rank matrix, the problems of unstable face similarity and poor style consistency in the synthetic image in the prior art are solved, and the image synthesis effect with high similarity and consistency is achieved, which improves training efficiency and model performance.
Patent Information
- Application Number
- CN202510388480.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
AI Technical Summary
When the prior art uses artificial intelligence models to synthesize images, it is difficult to achieve the effect of high similarity between users' faces and good style consistency. Especially when generating a large number of work photos of employees uniformly dressed, the similarity between the faces of the synthetic image and the original user is unstable and the style consistency is poor.
By combining portrait low-rank matrix and style low-rank matrix to construct an image synthesis model, the basic model is fine-tuned using the portrait fine-tuning dataset and style fine-tuning dataset respectively. The trained portrait low-rank matrix and style low-rank matrix ensure that the high similarity between faces and users and the consistency of styles in the synthetic image is maintained.
It achieves the effect of higher similarity between faces and users and better style consistency in the synthetic image, improves the stability and consistency of the image synthesis model, reduces workload and improves training efficiency.
Smart Images

Figure CN120387937A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to image synthesis, an image synthesis model training method, a medium, and a product. Background Art
[0002] With the development of image synthesis technology, especially in combination with artificial intelligence models, the image synthesis ability has been significantly improved.
[0003] In some work scenarios, there are unified requirements for user photos. For example, a certain company wants to take work photos of its employees. There are a large number of employees, and they are distributed all over the country. If they take photos with completely unified clothing, the workload is very large and it requires a lot of manpower and time. Therefore, the company wants to use an artificial intelligence model to help complete the work photos of employees with unified clothing. However, the synthesis effect is often not ideal. For example, it can synthesize images of the style specified by the user, but the similarity between the synthesized face and the original user is relatively low; or, when the similarity between the face and the original user is high, the consistency of the synthesized image styles of different users is poor. Summary of the Invention
[0004] The present disclosure provides an image synthesis method, an image synthesis model training method, a medium, and a product.
[0005] According to a first aspect of the present disclosure, an image synthesis method is provided. The method specifically includes: obtaining a user facial image to be synthesized and image style indication information provided by a user; inputting the user facial image and the image style indication information into an image synthesis model to generate a synthesized image with a specified style; wherein, the image synthesis model is constructed by using a first basic model in combination with a portrait low-rank matrix and a style low-rank matrix obtained through training; the portrait low-rank matrix and the style low-rank matrix are respectively obtained through training based on different basic models.
[0006] Based on the above content, it can be seen that by combining the first basic model with the pre-trained portrait low-rank matrix and style low-rank matrix, the image synthesis model has better image synthesis ability. When the portrait low-rank matrix and the style low-rank matrix are trained, they are not directly obtained through training using the first basic model, but first the first basic model is fine-tuned to obtain a second basic model and a third basic model. Then, the portrait low-rank matrix is obtained through training using the second basic model, and the style low-rank matrix is obtained through training using the third basic model. Finally, the first basic model and the two specifically trained low-rank matrices are combined together to obtain the required image synthesis model. Such an image synthesis model can achieve that the synthesized image maintains the consistency of different user styles and the difference of different user faces, and the similarity between the face in the synthesized image and the user is higher.
[0007] Before inputting the user's facial image into the image synthesis model according to at least one embodiment of the present disclosure, it further includes: extracting a face image from the user's facial image; performing downsampling processing on the face image to obtain a facial image matrix that matches the latent representation matrix in the image synthesis model, so as to input the facial image matrix into the image synthesis model.
[0008] According to at least one embodiment of the present disclosure, the training method of the portrait low-rank matrix includes: fine-tuning the first basic model using the portrait fine-tuning dataset to obtain a second basic model; training the portrait low-rank matrix using the portrait training samples and the second basic model.
[0009] According to at least one embodiment of the present disclosure, the training method of the style low-rank matrix includes: fine-tuning the first basic model using the style fine-tuning dataset to obtain a third basic model; training the style low-rank matrix using the style training samples and the third basic model.
[0010] According to a second aspect of the present disclosure, there is provided an image synthesis model training method. The method specifically includes: fine-tuning the first basic model using the portrait fine-tuning dataset and the style fine-tuning dataset to obtain a second basic model and a third basic model; training the portrait low-rank matrix using the portrait training samples and the second basic model; and training the style low-rank matrix using the style training samples and the third basic model; constructing an image synthesis model based on the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0011] According to at least one embodiment of the present disclosure, fine-tuning the first basic model using the portrait fine-tuning dataset and the style fine-tuning dataset to obtain a second basic model and a third basic model includes: obtaining the portrait fine-tuning dataset and the style fine-tuning dataset; fine-tuning the first basic model using the portrait fine-tuning dataset to obtain a second basic model; fine-tuning the first basic model using the style fine-tuning dataset to obtain a third basic model.
[0012] According to at least one embodiment of the present disclosure, training the portrait low-rank matrix using the portrait training samples and the second basic model; and training the style low-rank matrix using the style training samples and the third basic model includes: obtaining the portrait training samples and the style training samples; inputting the portrait training samples into the second basic model for training and outputting a portrait low-rank matrix compatible with the first basic model; inputting the style training samples into the third basic model for training and outputting a style low-rank matrix compatible with the first basic model; constructing an image synthesis model using the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0013] According to at least one embodiment of the present disclosure, after training using portrait training samples as input to a first basic model, the output is a portrait low-rank matrix, including: extracting a face from the portrait training samples to obtain a face image; performing downsampling processing on the face image to obtain a face mask matrix; optimizing the latent representation matrix in the first basic model using the face mask matrix; and training using the optimized first basic model to obtain a portrait low-rank matrix.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, such that the processor executes the method described in the first aspect or the second aspect of any one of the embodiments of the present disclosure.
[0015] According to a fourth aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, and when the execution instructions are executed by a processor, they are used to implement the method described in the first aspect or the second aspect of any one of the embodiments of the present disclosure.
[0016] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect or the second aspect of any one of the embodiments of the present disclosure. Description of the Drawings
[0017] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0018] Figure 1 It is a schematic flowchart of a model training method provided by the present disclosure.
[0019] Figure 2 It is a schematic diagram of the model training process illustrated by the present disclosure.
[0020] Figure 3 It is a schematic diagram of the portrait low-rank matrix training process provided by the present disclosure.
[0021] Figure 4 It is a schematic flowchart of an image synthesis method provided by the present disclosure.
[0022] Figure 5 It is a schematic block diagram of the structure of an image synthesis model training device according to an embodiment of the present disclosure.
[0023] Figure 6 It is a schematic block diagram of the structure of an image synthesis device according to an embodiment of the present disclosure.
[0024] Figure 7 Schematic diagram of the working photo generation process illustrated in this disclosure.
[0025] Figure 8 Block diagram showing the structure of an electronic device according to an embodiment of this disclosure. Detailed implementation manners
[0026] The following further elaborates on this disclosure with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the relevant content and do not limit this disclosure. Additionally, it should be noted that, for ease of description, only parts related to this disclosure are shown in the drawings.
[0027] It should be noted that, without conflict, the embodiments in this disclosure and the features in the embodiments can be combined with each other. The following will detail the technical solutions of this disclosure with reference to the drawings and in combination with the embodiments.
[0028] The field of portrait generation is one of the current hottest fields of Artificial Intelligence Generated Content (AIGC). By extracting the user's facial feature information and matching it with the corresponding style LoRA and the base model Diffusion, any style of user portrait can be produced. The current mainstream methods are divided into two types. One is LoRA-free, that is, directly extracting the user's facial features for inference without training, and the other is training the user portrait LoRA and performing inference while matching the style LoRA. However, the existing technical solutions have various problems. For example, the stability of the portrait is not good, that is, the similarity between the synthesized image and the provided face user is sometimes high and sometimes low, which is very unstable. There is also poor consistency in the generated image style. For example, when generating working photos, there are differences in the tie patterns among the working photos of employees in the same company. Therefore, there is an urgent need for a solution that can generate a synthetic image with a high similarity to the user's face and that conforms to a specified style.
[0029] For ease of description and to make the technical solutions of the detailed implementation manners of this disclosure easier to understand, before describing the image synthesis method implemented in this disclosure, the technical terms involved in the detailed implementation manners of this disclosure are explained as follows.
[0030] The LoRA model is a technology specifically used for fine-tuning large language models, which is referred to as a low-rank matrix in the following content. Different low-rank matrices with different functions need to be specifically trained using different training samples and base models. Its core idea is to introduce a low-rank matrix to reduce the number of parameters that need to be adjusted without changing the structure of the original large language model (that is, the base model mentioned later), thereby reducing the cost of fine-tuning while maintaining or nearly maintaining the original performance of the model.
[0031] Figure 1 The flowchart of a model training method provided by the present disclosure. As Figure 1 shown, the method includes steps 101 to 103. Among them, the method can be executed by an electronic device such as a server (local server or cloud server).
[0032] Specifically, Figure 1 the shown method includes: Step 101: Fine-tune the first basic model using the portrait fine-tuning dataset and the style fine-tuning dataset to obtain a second basic model and a third basic model.
[0033] When fine-tuning the first basic model, in order to meet different subsequent training requirements, different datasets are specifically selected for fine-tuning. Specifically, in order to meet the subsequent image synthesis requirements, many portrait images are collected to jointly form the portrait fine-tuning dataset. Then, the first basic model is fine-tuned using the portrait fine-tuning dataset to obtain a second basic model focused on image synthesis.
[0034] At the same time, in order to meet the subsequent synthesized image style requirements, many different style images are collected to jointly form the style fine-tuning dataset. Then, the first basic model is fine-tuned using the style fine-tuning dataset to obtain a third basic model focused on synthesizing images of a specified style.
[0035] The second basic model and the third basic model obtained in the above manner are more targeted. That is to say, in order to facilitate better subsequent training to obtain a portrait low-rank matrix (portrait LoRA) and a style low-rank matrix (style LoRA), dedicated basic models are respectively selected during training. So that the low-rank matrix obtained by training has stronger image synthesis ability after being combined with the first basic model.
[0036] Step 102: Train a portrait low-rank matrix using portrait training samples and the second basic model; and, train a style low-rank matrix using style training samples and the third basic model.
[0037] When training the portrait low-rank matrix and the style low-rank matrix, corresponding training samples need to be prepared. That is, portrait training samples and style training samples are respectively prepared. Specifically, a portrait low-rank matrix is trained using portrait training samples and the second basic model obtained by targeted fine-tuning. Since the second basic model is obtained by fine-tuning the first basic model using the portrait fine-tuning dataset, the portrait low-rank matrix obtained by training has better image synthesis ability than the portrait low-rank matrix directly trained using the first basic model (that is, the synthesized portrait is more similar to the user).
[0038] The style low-rank matrix is trained using the third basic model obtained by using the style training samples and targeted fine-tuning. Since the third basic model is obtained by fine-tuning the first basic model with the style fine-tuning data set, the trained style low-rank matrix has better style image synthesis ability than the style low-rank matrix directly trained using the first basic model (that is, the different user styles of the synthesized images are more consistent).
[0039] In addition, during training, the portrait low-rank matrix and the style low-rank matrix are trained separately and each uses a different basic model. This ensures that they do not interfere with each other during the training process, so that the two trained low-rank matrices are more targeted and their synthesis abilities are more balanced when synthesizing images.
[0040] Step 103: Construct an image synthesis model based on the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0041] Here, since both the portrait low-rank matrix and the style low-rank matrix are trained using the fine-tuning models of the first basic model, the portrait low-rank matrix and the style low-rank matrix have the same matrix dimensions and are both compatible with the first basic model. In addition, since the first basic model has not been fine-tuned at all, when the first basic model performs the image synthesis task, its overall synthesis ability is relatively balanced. The final image synthesis effect mainly depends on the two low-rank matrices. In other words, in the image synthesis model, the portrait low-rank matrix and the style low-rank matrix fully balance the synthesis effects of the portrait and the style during the image synthesis process, that is, making the face in the synthesized image more similar to the user and the styles between different users more consistent.
[0042] For example, as Figure 2 is a schematic diagram of the model training process illustrated in this disclosure. As can be seen from Figure 2 , first, the first basic model is respectively fine-tuned using the portrait fine-tuning data set and the style fine-tuning data set to obtain two basic models, namely the second basic model and the third basic model. On this basis, the portrait training samples and the style training samples can be further used for training to obtain the portrait low-rank matrix and the style low-rank matrix. Then, the portrait low-rank matrix and the style low-rank matrix can be inserted into the first basic model. During the training process, the portrait low-rank matrix and the style low-rank matrix are trained separately to avoid mutual influence. When fusing, the portrait low-rank matrix and the style low-rank matrix are inserted into the first basic model to ensure good compatibility and can be directly used without further debugging. Through the above solution, not only the training effect of the image synthesis model is improved, but also the training efficiency can be effectively improved.
[0043] In one or more embodiments of the present disclosure, fine-tuning a first base model using a portrait fine-tuning dataset and a style fine-tuning dataset to obtain a second base model and a third base model includes: obtaining a portrait fine-tuning dataset and a style fine-tuning dataset; fine-tuning the first base model using the portrait fine-tuning dataset to obtain the second base model; and fine-tuning the first base model using the style fine-tuning dataset to obtain the third base model.
[0044] To make the subsequent low-rank matrix training effect better, it is necessary to fine-tune the corresponding base model according to different needs. To facilitate the construction of the subsequent image synthesis model, the same base model can be used for fine-tuning. Here, it is assumed that the first base model is used for fine-tuning.
[0045] The portrait fine-tuning dataset mentioned here can, for example, contain a large number of real portrait images, covering different ethnic groups, genders, ages, expressions, lighting conditions, and backgrounds, etc., to ensure that the model can learn a wide range of portrait features. The portrait fine-tuning dataset may be sourced from public face databases, social media platforms, or professional photography libraries.
[0046] The style fine-tuning dataset mentioned here can be images with diverse styles and textures, used to train the model to capture and transform style features. The style fine-tuning dataset can include suits of different styles, ties of different styles, brooches of different styles, emblems of different companies, trademarks of different companies, etc., as well as corresponding style labels or descriptions.
[0047] Before fine-tuning, it may be necessary to preprocess the portrait fine-tuning dataset, such as image enhancement, size adjustment, normalization, etc., to ensure data consistency and the efficiency of model training. Using the portrait fine-tuning dataset to fine-tune the first base model, focusing on learning the fine features and generation ability of portraits. Specific loss functions and training strategies may be adopted during the fine-tuning process to optimize the performance of the model. During or after the fine-tuning process, an independent validation dataset is used to evaluate the performance of the model to ensure the accuracy and robustness of the model in portrait generation.
[0048] Next, fine-tune the third base model. Similar to the portrait fine-tuning dataset, the style fine-tuning dataset also needs to be preprocessed to ensure data consistency and the efficiency of model training. Using the style fine-tuning dataset to fine-tune the first base model, focusing on learning style features and transformation ability. Different loss functions and training strategies may be adopted during the fine-tuning process to adapt to the characteristics of the style transformation task. After the fine-tuning is completed, diverse style images are used to evaluate the style transformation ability of the model to ensure that the model can generate high-quality images in multiple styles.
[0049] It should be noted that if more synthesis requirements are needed during image synthesis, the fine-tuning of the basic model can be increased accordingly. Here, it is assumed to be the fourth basic model. Correspondingly, a fine-tuning dataset needs to be collected, and the first basic model is fine-tuned using this fine-tuning dataset to obtain the fourth basic model. In this way, the basic model is fine-tuned using different fine-tuning datasets according to different requirements, so that the low-rank matrix obtained by subsequent training has better image synthesis ability.
[0050] In addition, the above-mentioned solutions all fine-tune the first basic model using the dataset. Therefore, there is no need to prepare a large fine-tuning dataset, and the workload during the fine-tuning process is not very large. Through the above solutions, it is possible to improve the training effect of the model well without much workload.
[0051] In one or more embodiments of the present disclosure, a portrait low-rank matrix is trained using portrait training samples and a second basic model; and, a style low-rank matrix is trained using style training samples and a third basic model, including: obtaining portrait training samples and style training samples; inputting the portrait training samples into the first basic model for training and then outputting a portrait low-rank matrix compatible with the first basic model; inputting the style training samples into the second basic model for training and then outputting a style low-rank matrix compatible with the first basic model; constructing a synthesis model using the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0052] It should be noted that the portrait low-rank matrix can be called portrait LoRA, and the style low-rank matrix can be called style LoRA.
[0053] The portrait training samples mentioned here are samples that need to be labeled. The labeling may involve key point localization of the face, expression classification, pose estimation, etc. These information are crucial for training a LoRA that can capture and strengthen the key features in the portrait generation process. In addition, since LoRA aims to improve the generation ability of the model, these data may also include high-quality generation targets, such as high-resolution portrait images, and corresponding low-resolution or stylized input images.
[0054] The style training samples mentioned here are samples that need to be labeled. The labeling may involve style classification, style feature extraction, style conversion targets, etc. These information are crucial for training a LoRA that can capture and convert style features while keeping the content information unchanged. In addition, since LoRA aims to improve the style conversion ability of the model, these data may also include image pairs before and after style conversion, and corresponding style descriptions or labels.
[0055] During training, select a suitable base model, such as diffusion as the first base model, and then obtain the second base model through fine-tuning, which serves as the starting point for training the portrait LoRA. Input the preprocessed portrait training samples into the first base model for training. During the training process, the model optimizes its parameters through the backpropagation algorithm to minimize the loss function, thereby learning the general features of portraits. After training is completed, the model outputs a portrait low-rank matrix, namely the portrait LoRA. This matrix is a compact representation of the key information about portraits learned by the model during the training process. The portrait LoRA contains the encoding of portrait features by the model and can be used for subsequent portrait generation, recognition, or editing tasks.
[0056] Through the above training method, based on the second base model, use specific training strategies (such as contrastive learning, generative adversarial networks, etc.) and labeled portrait data to train the portrait LoRA. The goal of the portrait LoRA is to capture and strengthen the key features in the portrait generation process, enabling highly similar and high-quality portraits to be generated during inference. It should be noted that the portrait low-rank matrix (i.e., the portrait LoRA) obtained here should be well compatible with the first base model, that is, it has a matching parameter format, matrix dimension, and data type with the first base model.
[0057] Use the same first base model to perform fine-tuning to obtain the third base model, which serves as the starting point for training the style LoRA. Input the preprocessed style training samples into the second base model for training. During the training process, the model optimizes its parameters through the backpropagation algorithm to capture style features and minimize the loss with style labels or descriptions. After training is completed, the model outputs a style low-rank matrix, namely the style LoRA. This matrix is a compact representation of the key information about styles learned by the model during the training process. The style LoRA contains the encoding of style features by the model and can be used for subsequent style transfer, artistic creation, or image editing tasks. It should be noted that the portrait low-rank matrix (i.e., the portrait LoRA) obtained here should be well compatible with the first base model, that is, it has a matching parameter format, matrix dimension, and data type with the first base model.
[0058] Through the above training method, based on the third base model, use training methods related to style transfer and labeled style data to train the style LoRA. The goal of the style LoRA is to capture and transform the style features in the image while keeping the content information unchanged to achieve diverse style transfer effects.
[0059] In the synthesis model, the weights of the portrait LoRA and the style LoRA are combined with the weights of the first base model. This can be achieved through simple weight addition, weight interpolation, or more complex fusion strategies. During the subsequent application inference process, the weights of the portrait LoRA and the style LoRA can be adaptively adjusted according to the characteristics of the input image to achieve better generation effects. This can be achieved by introducing additional adaptive layers or attention mechanisms.
[0060] In the intermediate layer of the first base model, the features extracted by the portrait LoRA and the style LoRA can be fused. This can be achieved through methods such as feature concatenation, feature addition, or feature transformation to generate a synthetic image that contains both portrait and style features. After constructing the image synthesis model, a synthetic dataset containing portrait and style labels can be used for training and optimization. During the training process, appropriate loss functions and training strategies can be adopted to ensure that the model can learn the key features of the synthetic image and generate high-quality synthetic results.
[0061] In one or more embodiments of the present disclosure, training the first base model with portrait training samples and then outputting a portrait low-rank matrix includes: extracting the face from the portrait training samples to obtain a face image; performing downsampling processing on the face image to obtain a face mask matrix; optimizing the latent representation matrix in the second base model using the face mask matrix; and training with the optimized first base model to obtain a portrait low-rank matrix.
[0062] In practical applications, the face part is extracted from a large number of portrait training samples. This step usually involves using face detection algorithms such as Haar features, HOG+SVM, or deep learning models (such as MTCNN, FaceBoxes, etc.) to accurately locate and crop the face region, thereby obtaining a pure face image.
[0063] The extracted face image is subjected to downsampling processing to generate a face mask (i.e., a face mask matrix). The downsampling factor can be set according to specific requirements. For example, 8-fold downsampling is performed, which is consistent with the VAE downsampling factor in the fine-tuned diffusion model (i.e., the second base model). After downsampling, the foreground (i.e., the face part) in the face mask is marked as 1, and the background is marked as 0, forming a binary mask image.
[0064] The generated face mask is applied to the second base model to optimize the latent representation matrix. By combining the face mask with the latent representation of the model, the model can be guided to pay more attention to the feature representation of the face part, thereby optimizing the generation ability of the model.
[0065] Training is carried out using the optimized second basic model, and finally a portrait low-rank matrix is obtained. The low-rank matrix is a result of matrix decomposition, which can reduce the dimension and complexity of data while retaining key information. In portrait generation, the low-rank matrix may represent the main features and structural information of the face, and these information are crucial for generating high-quality portraits.
[0066] In actual operation, it may be necessary to multiply the face mask by the face representation in the latent space of the diffusion model to mask out information other than the face. This step ensures that the model only focuses on restoring the face part during the generation process, thereby improving the similarity between the generated portrait and the real portrait.
[0067] After training is completed, it is necessary to evaluate the generated portrait low-rank matrix. This can be done by comparing metrics such as the similarity and clarity between the generated portrait and the real portrait. According to the evaluation results, the model can be further adjusted and optimized to improve the quality of the generated portrait.
[0068] The obtained portrait low-rank matrix can be applied to various scenarios, such as portrait editing, image synthesis, face recognition, etc. In addition, this method can also be extended to other types of image generation tasks, such as animals, landscapes, etc., to explore more extensive application possibilities.
[0069] Based on the above scheme, by preprocessing the portrait training samples, downsampling, optimizing the model using the face mask, and training to obtain the portrait low-rank matrix and other steps, the quality and similarity of the generated portrait can be effectively improved.
[0070] For ease of understanding, the training process of the portrait low-rank matrix will be described below through specific embodiments.
[0071] As Figure 3 is a schematic diagram of the training process of the portrait low-rank matrix provided by the present disclosure. As can be seen from Figure 3 , the user takes a selfie in a noisy background. In this selfie image, it not only contains the user's facial image but also contains very complex background information.
[0072] Furthermore, the face is extracted from the user's selfie image to obtain the user's facial image that can be used later. The extracted user's facial image is downsampled, and the purpose is to generate a face mask. The downsampling factor can be set according to specific needs. For example, 8-fold downsampling is used to make the downsampling of the face image consistent with the downsampling factor of the VAE in the diffusion model. After downsampling, the foreground (i.e., the face part) in the face mask is marked as 1, and the background is marked as 0, forming a binary mask image.
[0073] Apply the generated face mask to the first basic model to optimize the latent representation matrix. Here, the "first basic model" may refer to a pre-trained generative model, such as a GAN, VAE, or diffusion model, etc. By combining the face mask with the latent representation of the model, the model can be guided to pay more attention to the feature representation of the face part, thereby optimizing the generative ability of the model.
[0074] Based on the same idea, the present disclosure also provides an image synthesis method. As Figure 4 is a schematic flowchart of an image synthesis method provided by the present disclosure. As can be seen from Figure 4 it, the method specifically includes the following steps: Step 401: Obtain the user's facial image to be synthesized provided by the user and image style indication information. Step 402: Input the user's facial image and the image style indication information into an image synthesis model to generate a synthesized image with a specified style; wherein, the image synthesis model is constructed by using the first basic model in combination with a portrait low-rank matrix and a style low-rank matrix obtained through training; the portrait low-rank matrix and the style low-rank matrix are obtained through training based on different basic models respectively.
[0075] When the user has an image synthesis requirement (for example, the user wants to generate a work photo that meets the standard requirements), the user can provide the server with the user's facial image taken by himself / herself. The user can also provide the required image style indication information in the form of a prompt.
[0076] Furthermore, the image synthesis model can perform an image synthesis task by using the user's facial image and the image style indication information.
[0077] The image style indication information mentioned here can be understood as the style of the image that the user finally wants to synthesize. For example, various image styles such as ID photos, portraits, and Chinese styles. More specifically, ID photos can be further divided into ID photos of Company A, ID photos of Company B, and so on. It should be noted that the more detailed the image style indication information provided by the user, the more in line with the user's requirements the final synthesized image style will be.
[0078] During the inference stage of image synthesis, although the second base model and the third base model are used to train the portrait low-rank matrix (i.e., portrait LoRA) and the style low-rank matrix (i.e., style LoRA) respectively, when actually generating portraits and performing style transfer, it returns to the original first base model and combines these two LoRAs. The key to this approach is that by training LoRAs on different fine-tuned models respectively, it can avoid task conflicts that may arise from training on the same base model (such as training portrait LoRA and style LoRA simultaneously based on the first base model), and at the same time, it can use the weight parameter adjustment of LoRA to fuse portrait and style features, ultimately generating portrait images that are both highly similar and conform to a specific style.
[0079] In summary, the portrait LoRA and the style LoRA are respectively trained on the second base model and the third base model fine-tuned with a specific dataset. This training strategy aims to improve the performance and adaptability of the model to meet the requirements of complex and diverse generation tasks.
[0080] In one or more embodiments of the present disclosure, before inputting the user's facial image into the synthesis model, it further includes: extracting a face image from the user's facial image; performing downsampling processing on the face image to obtain a facial image matrix that matches the latent representation matrix in the synthesis model, so as to input the facial image matrix into the synthesis model.
[0081] For ease of understanding, the following will take the combination of portrait LoRA and the first base model Diffusion as an example for elaboration.
[0082] The Diffusion model is a generative model that generates data by learning the inverse diffusion process of the data distribution. In the field of image generation, the Diffusion model can gradually generate high-quality images from random noise. This model usually includes an encoder (encoding an image into a latent representation) and a decoder (decoding an image from the latent representation).
[0083] The Latent space refers to the latent representation space of the Diffusion model, that is, the vector space encoded by the encoder from the image. In this space, similar images will be mapped to nearby points, while different images will be mapped to farther points. The Latent space is the key to image generation and editing in the Diffusion model.
[0084] The Portrait LoRA is used to optimize a face generation or recognition task in combination with a Diffusion model (i.e., the second base model). Specifically, the face is extracted and combined with the representation in the Latent space of the Diffusion model. This combination is achieved by downsampling the face mask and multiplying it with the face representation in the Latent space, thus masking out all background information and enabling the model to focus more on restoring the face part.
[0085] Specifically, the face region is extracted from the user-provided user facial image. Face mask downsampling: The extracted face mask is downsampled by 8 times to match the downsampling factor of the VAE (Variational Autoencoder) of the Diffusion model. The downsampled face mask is multiplied with the face representation in the Latent space of the Diffusion model to mask out the background information.
[0086] In this way, the Portrait LoRA and the Diffusion model are used in combination to optimize the face generation or recognition task. The Portrait LoRA provides the ability to efficiently fine-tune the model, while the Diffusion model provides powerful image generation capabilities. This combination enables the model to more accurately restore or recognize the face while keeping the background information unaffected.
[0087] Next, the generation process of the portrait low-rank matrix and the style low-rank matrix will be described through specific embodiments.
[0088] The training method of the portrait low-rank matrix includes: fine-tuning the first base model using the portrait fine-tuning dataset to obtain the second base model; training the portrait low-rank matrix using the portrait training samples and the second base model.
[0089] The training method of the style low-rank matrix includes: fine-tuning the first base model using the style fine-tuning dataset to obtain the third base model; training the style low-rank matrix using the style training samples and the third base model.
[0090] When training the image synthesis model, targeted fine-tuning is first performed based on the same first base model. When adjusting, different fine-tuning datasets need to be prepared respectively.
[0091] Specifically, prepare the portrait fine-tuning dataset. The portrait data in this portrait fine-tuning dataset are portrait images provided with user authorization. A fine-tuning dataset containing high-quality and diverse portraits. This fine-tuning dataset should be similar to the portrait training samples, but may be more focused or have specific styles, expressions, etc. features, or facial images taken from various angles, aiming to further refine the outstanding capabilities of the base model in portrait generation requirements.
[0092] When fine-tuning the first basic model, first, select a pre-trained first basic model, which may be a powerful generative model such as a GAN, VAE, diffusion model, etc. This model already has the ability to generate high-quality images.
[0093] When fine-tuning the first basic model, input the portrait fine-tuning dataset into the first basic model and fine-tune the parameters of the first basic model through the backpropagation algorithm. The goal of fine-tuning is to make the first basic model better capture the subtle features and styles of portraits while maintaining its ability to generate high-quality images. After fine-tuning, a new basic model, namely the second basic model, is obtained. This model is more accurate in portrait feature representation than the original basic model.
[0094] After obtaining the second basic model, further training can be carried out using portrait training samples and the second basic model, with the goal of obtaining a portrait low-rank matrix, that is, portrait LoRA. This matrix contains the key information of the model in portrait feature representation and can be used for subsequent portrait generation or editing tasks.
[0095] When training the style low-rank matrix, collect a fine-tuning dataset containing various styles and textures. This dataset should cover a wide range of style types so that the model can learn rich style features.
[0096] It should be noted that the basic model used for the style low-rank matrix is also obtained by fine-tuning the first basic model. Select a pre-trained basic model as the starting point.
[0097] Input the style fine-tuning dataset into the first basic model and fine-tune the parameters of the first basic model through the backpropagation algorithm. The goal of fine-tuning is to enable the first basic model to capture and represent style features while maintaining its image generation ability.
[0098] After fine-tuning, a new model, that is, the third basic model, is obtained. This third basic model is more excellent in style feature representation than the original basic model.
[0099] Furthermore, use style training samples and the third basic model for training, with the goal of obtaining a style low-rank matrix, that is, style LoRA. This matrix contains the key information of the model in style feature representation and can be used for subsequent style transfer or artistic creation tasks.
[0100] Based on any of the above embodiments, the present disclosure also provides an image synthesis model training device. Figure 5 The structural schematic block diagram of the image synthesis model training device according to an embodiment of the present disclosure. As Figure 5As shown in the figure, the image synthesis model training device includes: a fine-tuning module 51, configured to fine-tune a first basic model by using a portrait fine-tuning data set and a style fine-tuning data set to obtain a second basic model and a third basic model.
[0101] A training module 52, configured to train a portrait low-rank matrix by using portrait training samples and the second basic model; and, train a style low-rank matrix by using style training samples and the third basic model.
[0102] A construction module 53, configured to construct an image synthesis model based on the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0103] The fine-tuning module 51 is further configured to obtain a portrait fine-tuning data set and a style fine-tuning data set; fine-tune the first basic model by using the portrait fine-tuning data set to obtain a second basic model; fine-tune the first basic model by using the style fine-tuning data set to obtain a third basic model.
[0104] The training module 52 is further configured to obtain portrait training samples and style training samples; input the portrait training samples into the second basic model for training and then output a portrait low-rank matrix compatible with the first basic model; input the style training samples into the third basic model for training and then output a style low-rank matrix compatible with the first basic model; construct an image synthesis model by using the first basic model, the portrait low-rank matrix, and the style low-rank matrix.
[0105] The training module 52 is further configured to extract a face from the portrait training samples to obtain a face image; perform downsampling processing on the face image to obtain a face mask matrix; optimize a latent representation matrix in the first basic model by using the face mask matrix; and train a portrait low-rank matrix by using the optimized first basic model.
[0106] Based on any of the above embodiments, the present disclosure further provides an image synthesis device. Figure 6 It is a structural schematic block diagram of an image synthesis device according to an embodiment of the present disclosure. As Figure 6 shown, the image synthesis device includes: an acquisition module 61, configured to acquire a user facial image to be synthesized provided by a user and image style indication information.
[0107] An image synthesis module 62, configured to input the user facial image and the image style indication information into an image synthesis model to generate a synthesized image with a specified style. Wherein, the image synthesis model is constructed by using the first basic model in combination with a portrait low-rank matrix and a style low-rank matrix obtained through training; the portrait low-rank matrix and the style low-rank matrix are respectively obtained through training based on different basic models.
[0108] Optionally, it further includes a preprocessing module 63 for extracting a face image from the user face image; performing downsampling processing on the face image to obtain a face image matrix matching the latent representation matrix in the image synthesis model, so as to input the face image matrix into the image synthesis model.
[0109] Optionally, it further includes a training module 64 for fine-tuning the first basic model using a portrait fine-tuning data set to obtain a second basic model; training the portrait low-rank matrix using portrait training samples and the second basic model.
[0110] The training module 64 is used to fine-tune the first basic model using a style fine-tuning data set to obtain a third basic model; training the style low-rank matrix using style training samples and the third basic model.
[0111] For ease of understanding, the following will be elaborated through specific embodiments. As Figure 7 is a schematic diagram of the work photo generation process illustrated in this disclosure. As can be seen Figure 7 from it, after collecting the image containing the user's face information provided by the user, operations such as quality filtering, skin beautification and flaw removal, and head extraction will be performed to obtain a more accurate and feature-prominent user face image. Then, the obtained image is given to the image synthesis model to perform the image synthesis task.
[0112] The image synthesis model mentioned here can be trained using portrait training samples and style training samples. Specifically, first select a basic model, which is assumed to be the first basic model here. Then, perform targeted fine-tuning on the first basic model to obtain a second basic model for training the portrait low-rank matrix and a third basic model for training the style low-rank matrix. It can be seen here that different basic models are used for training during training, so as to avoid the mutual influence of the two low-rank matrices when training in the same basic model. In other words, the image processing capabilities of the portrait LoRA and style LoRA trained in the above manner are better.
[0113] After completing the training, the portrait low-rank matrix and the style low-rank matrix are simultaneously inserted into the first basic model (Stable diffusion). These two low-rank matrices have a better compatibility effect with the first basic model. After insertion, there is no need for cumbersome debugging, and a usable image synthesis model can be directly obtained.
[0114] Furthermore, the previously obtained user face image and the image style indication information proposed by the user are input into the image synthesis model. To improve the synthesis effect, pose control, quality filtering, similarity filtering, face fusion, etc. can also be performed on the synthesized image.
[0115] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0116] The execution subject of the image synthesis method and the image synthesis model training method in the specific embodiments of the present disclosure can be an electronic device such as a server (including a local server or a cloud server).
[0117] Therefore, based on any of the above embodiments, the present disclosure further provides an electronic device, which can execute the image synthesis method and the image synthesis model training method of any of the above embodiments described in the present disclosure.
[0118] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0119] Figure 8 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure.
[0120] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0121] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one connecting line is used in this figure, but it does not mean that there is only one bus or one type of bus.
[0122] The present disclosure also provides a readable storage medium storing a computer program which, when executed by a processor, is used to implement the above method. A "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0123] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0124] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0125] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing method devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing method devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing method device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing method device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0129] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0130] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0131] Those skilled in the art should understand that the above embodiments are merely for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. An image synthesis method, characterized in that, The method includes: Obtaining a user facial image to be synthesized and image style indication information provided by a user; Inputting the user facial image and the image style indication information into an image synthesis model to generate a synthesized image with a specified style; Wherein, the image synthesis model is constructed by using a first basic model in combination with a portrait low-rank matrix and a style low-rank matrix obtained through training; the portrait low-rank matrix and the style low-rank matrix are obtained through training based on different basic models respectively.
2. The method according to claim 1, characterized in that, Before inputting the user facial image into the image synthesis model, it further includes: Extracting a face image from the user facial image; Performing downsampling processing on the face image to obtain a facial image matrix matching the latent representation matrix in the image synthesis model, so as to input the facial image matrix into the image synthesis model.
3. The method according to claim 1, characterized in that The training method of the portrait low-rank matrix includes: Fine-tuning the first basic model by using a portrait fine-tuning data set to obtain a second basic model; Training the portrait low-rank matrix by using portrait training samples and the second basic model.
4. The method according to claim 1, characterized in that, The training method of the style low-rank matrix includes: Fine-tuning the first basic model by using a style fine-tuning data set to obtain a third basic model; Training the style low-rank matrix by using style training samples and the third basic model.
5. A method for training an image synthesis model, characterized in that, The method includes: Fine-tuning a first basic model by using a portrait fine-tuning data set and a style fine-tuning data set to obtain a second basic model and a third basic model; Training a portrait low-rank matrix by using portrait training samples and the second basic model; and training a style low-rank matrix by using style training samples and the third basic model; Constructing an image synthesis model based on the first basic model, the portrait low-rank matrix and the style low-rank matrix.
6. The method according to claim 5, characterized in that, The fine-tuning of the first basic model by using a portrait fine-tuning data set and a style fine-tuning data set to obtain a second basic model and a third basic model includes: Obtaining a portrait fine-tuning data set and a style fine-tuning data set; Fine-tuning the first basic model by using the portrait fine-tuning data set to obtain a second basic model; Fine-tuning the first basic model by using the style fine-tuning data set to obtain a third basic model.
7. The method according to claim 5, wherein The training of the portrait low-rank matrix by using portrait training samples and the second basic model; And the training of the style low-rank matrix by using style training samples and the third basic model includes: Obtaining portrait training samples and style training samples; Inputting the portrait training samples into the second basic model for training and outputting the portrait low-rank matrix compatible with the first basic model; Inputting the style training samples into the third basic model for training and outputting the style low-rank matrix compatible with the first basic model.
8. The method according to claim 7, characterized in that The training of the portrait low-rank matrix by inputting the portrait training samples into the first basic model and outputting the portrait low-rank matrix includes: Performing face extraction on the portrait training samples to obtain a face image; Performing downsampling processing on the face image to obtain a face mask matrix; Optimizing the latent representation matrix in the first basic model by using the face mask matrix; The portrait low-rank matrix is obtained by training with the optimized first basic model.
9. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.