Personalized image generation method and system based on diffusion model rapid optimization

Through the personalized image generation method based on diffusion model, the multimodal big model and hidden space diffusion model are used, combined with the encoder and the celebrity condition regularization loss function, the problem of insufficient personalized generation ability of the existing model when processing specific individual images is solved, and high-quality and diverse personalized image generation is achieved.

CN120198538APending Publication Date: 2025-06-24TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102784.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing text-based image generation models are difficult to accurately generate personalized content through a single text description when processing images of specific individuals, resulting in limited applicability when personalizing needs.

Method used

A personalized image generation method based on rapid optimization of diffusion model is proposed. By obtaining the text feature vector of text input instructions, using multimodal large model and hidden space diffusion model training, combining encoder and replacement text features to generate images containing specific characters, and optimization of the results are generated through the regularization loss function of celebrity condition.

Benefits of technology

Given any face image and text instructions, a high-quality image with both identity characteristics and text description is achieved, which improves the diversity and accuracy of generated images and meets the generation needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198538A_ABST
    Figure CN120198538A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized image generation method and system based on diffusion model rapid optimization, and the method comprises the steps: obtaining a text feature vector corresponding to a text input instruction, coding a preset figure image into a text space, and obtaining a new text feature, and replacing words in the text feature and the text feature vector to generate an image containing a specific character. According to different text input instructions, encoding an image containing a specific character to obtain an encoded character image; re-coding the coded figure image based on a preset name of a celebrity to obtain a re-coded figure image; and constructing a celebrity condition regularization loss function based on the following capability of the celebrity name to the text input instruction so as to optimize the recoding figure image and obtain a final personalized figure image. According to the invention, under the condition that any face image and any character instruction are given, the image which accords with the character instruction and is based on the given face identity information can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of next-generation Internet application security, multi-modal large model security evaluation, and computer vision generation technology, and particularly relates to a personalized image generation method and system based on rapid optimization of a diffusion model. Background Art

[0002] The research background of the optimization method for specific characters based on the Text to Image (T2I) model is rich and diverse, covering the potential and limitations of current T2I models after training on large-scale datasets. Models such as GLIDE, DALL-E2, Imagen, and StableDiffusion (SD) have demonstrated their ability to generate high-quality images according to diverse text prompts, and are widely used in fields such as art creation, virtual product design, and other visual content generation areas. At the same time, they also pose severe security risks to Internet applications based on image authentication. However, although these models perform excellently in image generation, they still have obvious bottlenecks when dealing with specific concepts, especially when processing specific identity features such as a person's face. When faced with images of specific individuals, these models often have difficulty precisely generating personalized content through a single text description. For example, when users hope that the T2I model generates images with specific facial features of family members, friends, or themselves, although detailed text prompts are provided, existing models still have difficulty accurately reflecting these personalized information. This limits the applicability of these models in dealing with personalized needs and hinders their wide popularization in the field of user-generated content. To address this challenge, researchers have begun to focus on how to embed personalized concepts into existing large-scale T2I models without retraining the models. Traditional solutions, such as introducing conversion modules while freezing the models and adapting to new inputs by fine-tuning a small number of new concepts. However, these methods still have many problems, such as forgetting existing knowledge and difficulty in balancing new and old concepts. In parallel, another important application of the T2I model is to generate scene images related to individual identities through natural language descriptions, such as generating unique images for specific celebrities, users on social media, or artists. However, the current models' learning of identity features often overly relies on large-scale image-text pairs in the training data and lacks the ability to customize generation for specific individual features. Therefore, researchers have proposed different solutions, attempting to improve the performance of the models in generating images of specific individuals through a fast adjustment framework that combines identity retention and instruction execution capabilities. Generally speaking, the research on T2I models optimized for specific characters shows broad prospects in personalized generation. Researchers are committed to improving the models' ability to process specific individual identity information without compromising their image generation ability, promoting the technological progress of personalized visual content generation. This not only provides a new direction for personalized image generation but also brings more possibilities for future multi-field applications. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems in the related art to a certain extent.

[0004] The present invention proposes a personalized image generation method based on rapid optimization of diffusion models. Taking task image data as an example, it aims to achieve efficient and high-quality generation of specific person images based on arbitrary text instructions.

[0005] Another object of the present invention is to propose a personalized image generation system based on rapid optimization of diffusion models.

[0006] To achieve the above object, on the one hand, the present invention proposes a personalized image generation method based on rapid optimization of diffusion models, including:

[0007] Obtain the text feature vector corresponding to the text input instruction, and train a multi-modal large model through a diffusion model and a latent space diffusion model;

[0008] Use an encoder to encode a preset person image into the text space to obtain a new text feature, and generate an image containing a specific person by replacing the words in the text feature and the text feature vector;

[0009] Use the multi-modal large model to encode the image containing the specific person according to different text input instructions to obtain an encoded person image;

[0010] Re-encode the encoded person image based on a preset celebrity name to obtain a re-encoded person image;

[0011] Construct a celebrity-conditioned regularization loss function based on the ability of the celebrity name to follow the text input instruction, and optimize the re-encoded person image to obtain the final personalized person image.

[0012] The personalized image generation method based on rapid optimization of diffusion models in the embodiments of the present invention may also have the following additional technical features:

[0013] In an embodiment of the present invention, obtaining the text feature vector corresponding to the text input instruction includes:

[0014] Define the text input instruction of natural language text;

[0015] Segment the text input instruction into multiple words according to a preset vocabulary, and obtain the feature representations of each word through dictionary query;

[0016] Use an encoder to encode each word to obtain the text feature vector corresponding to the text input instruction.

[0017] In an embodiment of the present invention, the training objective of generating an image containing a specific person by replacing the words in the new text feature and the text feature vector is as follows:

[0018]

[0019] Among them, c(y, v) represents replacing the person words in y with the representation of the preset person image I, where v * is the text feature.

[0020] In an embodiment of the present invention, based on the multimodal large model M, the text input instruction y and the preset person image I, the trained multimodal large model is used to encode the person image I to obtain the encoding:

[0021] v = M(y′, I)

[0022] Among them, y′ represents the concatenation result of the facial details describing this person and the original text input instruction y.

[0023] In an embodiment of the present invention, the text encoder is used to encode the celebrity name to obtain the mean and standard deviation of the celebrity encoding: μ celeb and σ celeb ; Based on the encoded person image v = M(y′, I) obtained, after normalization, re-encoding is performed:

[0024]

[0025] In an embodiment of the present invention, constructing a celebrity-conditioned regularization loss function based on the ability of the celebrity name to follow the text input instruction to optimize the re-encoded person image to obtain the final personalized person image further includes:

[0026] Based on the celebrity name and the text input instruction, a series of reference image sets R are generated using the trained multimodal large model;

[0027] Constructing a celebrity-conditioned regularization loss function:

[0028]

[0029] Among them, is a reference frame randomly selected from the reference image set R according to the text input instruction y, and M represents the mask of the area outside the face region in

[0030] In an embodiment of the present invention, the feature vector c(y) corresponding to the text input instruction y is obtained:

[0031] c(y) = [E(v1, v2,..., v n )]

[0032] Among them, E represents the pre-trained text encoder, and n represents the number of words.

[0033] In one embodiment of the present invention, a text input instruction is input into a multi-modal large model to output an image x that conforms to the text semantic information; the method further includes:

[0034] Define the image compression encoder and decoder as ε and where the encoder compresses the input image x into the latent space to obtain z = ε(x), and the decoder maps the eigenvalue in the latent space back to the image space to obtain

[0035] In one embodiment of the present invention, through the compression of the image x, it is modeled based on the diffusion model:

[0036]

[0037] where t represents the time tag of the diffusion model, and ε is the noise randomly sampled at the current moment.

[0038] To achieve the above object, on the other hand, the present invention proposes a personalized image generation system based on rapid optimization of the diffusion model, including:

[0039] A model training module, configured to obtain the text feature vector corresponding to the text input instruction, and train a multi-modal large model through the diffusion model and the latent space diffusion model;

[0040] A specific person image encoding module, configured to use the encoder to encode a preset person image into the text space to obtain a new text feature, and generate an image containing the specific person by replacing the words in the text feature and the text feature vector;

[0041] An image encoding module based on text description conditions, configured to use the multi-modal large model to encode the image containing the specific person according to different text input instructions to obtain an encoded person image;

[0042] A word vector encoding module based on the celebrity name, configured to re-encode the encoded person image based on the preset celebrity name to obtain a re-encoded person image;

[0043] A person image optimization module, configured to construct a celebrity condition regularization loss function based on the ability of the celebrity name to follow the text input instruction, so as to optimize the re-encoded person image to obtain the final personalized person image.

[0044] The personalized image generation method and system based on rapid optimization of diffusion models according to the embodiments of the present invention can generate image content with both identity features and text descriptions when providing any face image and text instructions. By combining the optimization of identity information retention and text instruction compliance, this method greatly improves the quality of the generated images. Specifically, the present invention learns a set of exclusive embedding representations based on the given text prompt, thereby enhancing the retention effect of identity information during the generation process. Compared with the prior art, the present invention shows stronger adaptability in dealing with diverse text instructions and can generate images that conform to both identity features and respond to multiple prompts. This adaptability enables the generated images to retain the unique identity features of the person under different instructions and can flexibly change to meet the user's text description requirements. In addition, a celebrity conditional regularization loss mechanism is innovatively introduced to naturally drive the generation of person-centered images using celebrity names. This mechanism not only improves the local consistency of the generated images but also enhances their diversity and accuracy when facing different prompts. Through this mechanism, the generated images can maintain identity consistency under different text instructions and also show richer image details and style variations, thus meeting the generation requirements of different scenarios.

[0045] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0047] Figure 1 is a flowchart of a personalized image generation method based on rapid optimization of diffusion models according to an embodiment of the present invention;

[0048] Figure 2 is a structural diagram of a personalized image generation system based on rapid optimization of diffusion models according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0050] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] The following describes a personalized image generation method and system based on rapid optimization of a diffusion model according to an embodiment of the present invention with reference to the accompanying drawings.

[0052] Figure 1 is a flowchart of a personalized image generation method based on rapid optimization of a diffusion model according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0053] S1, obtaining a text feature vector corresponding to a text input instruction, and training a multi-modal large model through a diffusion model and a latent space diffusion model;

[0054] S2, using an encoder to encode a preset person image into a text space to obtain a new text feature, and generating an image containing a specific person by replacing words in the text feature and the text feature vector;

[0055] S3, using the multi-modal large model to encode the image containing a specific person according to different text input instructions to obtain an encoded person image;

[0056] S4, re-encoding the encoded person image based on a preset celebrity name to obtain a re-encoded person image;

[0057] S5, constructing a celebrity-conditioned regularization loss function based on the ability of the celebrity name to follow the text input instruction to optimize the re-encoded person image to obtain a final personalized person image.

[0058] It can be understood that the core objective of the present invention is to generate an image with specific person identity characteristics and consistent with the user's text description, and the adopted solution is based on the existing text-to-image (T2I) large model. In the specific implementation process, multiple large models can be selected. In the process of elaborating the details of the present invention, the present invention takes Stable Diffusion (SD) as an example for development.

[0059] The training of the T2I large model is completed based on a large dataset of image-text pairs. In the present invention, the disclosed open-source model weights are used as the base model. Taking Stable Diffusion (SD) as an example of the base model used in the present invention, this model is constructed through the training methods of diffusion models and latent space diffusion models to achieve the image generation task.

[0060] In an embodiment of the present invention, the input signal of the SD model is input into the model in the form of natural language text, and this text input signal is defined as y in the present invention. According to a pre-set vocabulary, the input signal y is first segmented into multiple words, and the feature representations v of each word are obtained through dictionary lookup i , where i represents the i-th segmented word. Subsequently, an encoder is used to encode each word to obtain the feature vector c(y) corresponding to the input signal y as follows:

[0061] c(y) = [E(v1, v2, …, v n )]

[0062] where E represents the pre-trained text encoder and n represents the number of words.

[0063] The training of the T2I large model is achieved by the method of "text reconstructing image". The present invention first defines the image compression encoder and decoder as ε and respectively, where the encoder can compress the input image x into the latent space to obtain z = ε(x), and at the same time the decoder can map the feature values in the latent space back to the image space to obtain Through the compression of the image, the training of the T2I large model can be more efficiently modeled based on the Diffusion Model (Denoising Diffusion Probabilistic Model) as follows:

[0064]

[0065] where t represents the time label of the diffusion model and ε is the noise randomly sampled at the current moment. The goal of model training is to predict the noise at the current moment based on the parameterized model ε θ to achieve the ability of gradually removing noise.

[0066] Furthermore, based on the above steps, there is already a way to input the text y and obtain the image x that conforms to the text semantic information. Here, the present invention elaborates on how to generate a specific person's image based on this, in combination with the specific person's image I.

[0067] Among them, the present invention first introduces the general method of using an encoder to encode the person's image I into the text space to obtain a new "word" v* This "word" shares the same feature space with other text words, but does not overlap with other words and only represents a given target person. Thus, generating an image of this person can be achieved by replacing a specific word in v * and the text feature vector c(y). For example, given y as "a person skiing", an image of a specific person skiing can be generated by replacing the text feature corresponding to "a person" with v * . The training objective of this scheme is as follows:

[0068]

[0069] where c(y, v) represents the representation of replacing the person word in y with a specific person image I. Usually, y uses descriptions such as "a photo of a person", "a person's headshot in the middle of the image", etc.

[0070] Based on the encoding of the specific person image described previously, the present invention proposes to utilize the text understanding ability of a pre-trained multi-modal large model and combine different input text instructions to encode the person image, so that the result is more in line with the requirements of different input text instructions. Based on a multi-modal large model M, text instruction y, and person image I, the present invention first encodes the person image I using the ability of the pre-trained multi-modal large model: v = M(y′, I). Where y′ represents the concatenation result of "describe the facial details of this person" and the original text instruction y.

[0071] It has been experimentally found that using the name of a celebrity can directly generate an image centered on that person. Therefore, the present invention proposes to use the name of a celebrity as a feature variation basis, so that the encoded v has stronger editability. The present invention collects the names of some celebrities and encodes them using a text encoder to obtain a celebrity encoding mean and standard deviation: μ celeb and σ celeb . Based on the previously obtained person image encoding v = M(y′, I), after normalizing it, it is re-encoded:

[0072]

[0073] Furthermore, based on the previous step, the present invention can already use a specific image and a pre-trained T2I model for rapid optimization to generate an image of a specific person. However, due to the lack of rich text-image pairs and detailed text descriptions, the results obtained by such training usually have problems such as a single scene and weakened instruction-following ability. Here, based on the instruction-following ability of the celebrity name for text input, the present invention proposes a celebrity-conditioned regularization loss function.

[0074] The present invention first uses the celebrity names collected previously as a basis, combines a rich dataset of text instructions, and uses a pre-trained T2I model to generate a series of reference image sets R.

[0075] Ideally, given any person image I, the optimized model should be able to generate a scene that conforms to the reference image set R and a face region that conforms to the given image I. Based on this idea, the present invention proposes the following celebrity conditional regularization loss function:

[0076]

[0077] where is a reference frame randomly selected from the reference image set R according to the text instruction y, and M represents a mask for the region outside the face region in . In this way, the model can learn to reconstruct the face information in the image I while maintaining the semantics of the text instruction y.

[0078] In summary, the present invention is mainly used for personalized person image generation guided by text instructions. The invention can generate an image that conforms to the text instruction and is based on the given face identity information given any face image and any text instruction. The invention optimizes from two aspects of identity information retention and instruction execution, and realizes a fast optimization scheme with better comprehensive performance. Specifically, the method aims to improve the balance between identity information and text instruction conditions in the prior art, and learn a set of exclusive embedding representations under the condition of a given prompt word, so as to enhance the retention effect of identity information in the generation process. Compared with traditional methods, the invention shows stronger adaptability in the effects of generating different prompt word scenarios, making the generated images not only retain the user's identity characteristics but also respond to different generation instructions. In addition, the method innovatively introduces a celebrity conditional regularization loss function. This loss function is based on the findings in existing research, which believes that celebrity names can naturally drive person-centered generation. By using the celebrity conditional regularization loss function, the present invention further strengthens the local consistency of the generation results under different prompt conditions, enabling the generated images to better follow the generation instructions while maintaining the identity, and enhancing the diversity and accuracy of the generated images. Finally, the present invention makes a comprehensive comparison with existing advanced schemes for the task of personalized person image generation guided by text instructions. In terms of visual effects, the scheme of the present invention can generate richer and more diverse image contents and maintain the consistency of person identity information. In terms of objective evaluation indicators, the scheme of the present invention gives the optimal comprehensive evaluation indicators. The experimental results show the stronger robustness and generation quality of the invention, making it very suitable for application to content creation platforms in Internet platforms to help content creators quickly generate image contents with specific person information.

[0079] The personalized image generation method based on rapid optimization of diffusion model according to the embodiments of the present invention can generate image content with both identity features and text descriptions when providing any face image and text instruction. By combining the optimization of identity information retention and text instruction compliance, this method greatly improves the quality of the generated images. Compared with the prior art, the present invention not only shows higher diversity and fineness in visual effects, but also achieves the best comprehensive performance in quantitative evaluation. This makes the present invention have broad application potential on content creation platforms, especially suitable for personalized character image generation tasks. The present invention provides a more convenient way for content creators to quickly generate images with specific character identity features and meeting user text descriptions, thus greatly improving the efficiency and quality of content creation.

[0080] To implement the above embodiments, as Figure 2 shown, the personalized image generation system 10 based on rapid optimization of diffusion model is further provided in this embodiment, including:

[0081] A model training module 100, configured to obtain a text feature vector corresponding to a text input instruction, and train a multi-modal large model through a diffusion model and a latent space diffusion model;

[0082] A specific person image encoding module 200, configured to encode a preset person image into the text space by using an encoder to obtain a new text feature, and generate an image containing the specific person by replacing words in the text feature and the text feature vector;

[0083] An image encoding module 300 based on text description conditions, configured to encode the image containing the specific person according to different text input instructions by using the multi-modal large model to obtain an encoded person image;

[0084] A word vector encoding module 400 based on famous person names, configured to re-encode the encoded person image based on a preset famous person name to obtain a re-encoded person image;

[0085] A person image optimization module 500, configured to construct a famous person condition regularization loss function based on the ability of the famous person name to follow the text input instruction, so as to optimize the re-encoded person image to obtain a final personalized person image.

[0086] The personalized image generation system based on rapid optimization of diffusion models according to the embodiments of the present invention can generate image content with both identity features and text descriptions when providing any face image and text instructions. By combining the optimization of identity information retention and text instruction compliance, this method greatly improves the quality of the generated images. Compared with the prior art, the present invention not only shows higher diversity and fineness in visual effects, but also achieves the best comprehensive performance in quantitative evaluation. This makes the present invention have broad application potential on content creation platforms, especially suitable for personalized character image generation tasks. The present invention provides a more convenient way for content creators to quickly generate images with specific character identity features and meeting user text descriptions, thus greatly improving the efficiency and quality of content creation.

[0087] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0088] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

Claims

1. A personalized image generation method based on rapid optimization of diffusion model, characterized in that: include: Obtain the text feature vector corresponding to the text input instruction, and obtain a multimodal large model through diffusion model and latent space diffusion model training; Using an encoder to encode a preset character image into a text space to obtain a new text feature, and generating an image containing a specific character by replacing the text feature and words in the text feature vector; Using the multimodal large model to encode an image containing a specific person according to different text input instructions to obtain an encoded person image; Re-encoding the encoded character image based on a preset celebrity name to obtain a re-encoded character image; A celebrity-conditional regularized loss function is constructed based on the ability of celebrity names to follow text input instructions to optimize the re-encoded character image to obtain the final personalized character image.

2. The method according to claim 1, characterized in that Get the text feature vector corresponding to the text input instruction, including: Define text input instructions for natural language characters; The text input instruction is divided into multiple words according to a pre-set vocabulary, and the feature representation of each word is obtained by dictionary query; The encoder is used to encode each word to obtain the text feature vector corresponding to the text input instruction.

3. The method according to claim 1, characterized in that The training objective for generating images containing specific persons by replacing new text features and words in the text feature vector is as follows: Among them, c(y,v) represents the replacement of the character word in y with the representation of the preset character image I, and v * is a text feature.

4. The method according to claim 1, characterized in that Based on the multimodal large model M, the text input instruction y and the preset character image I, the character image I is encoded using the ability of the trained multimodal large model to obtain the encoding: v=M(y ′ ,I) Among them, y ′ Represents the concatenation of the facial details describing this character and the original text input instruction y.

5. The method according to claim 1, characterized in that Use the text encoder to encode the celebrity names to get the celebrity encoding mean and standard deviation: μ celeb and σ celeb ; Based on the obtained coded character image v = M (y ′ ,I), normalize and then re-encode:

6. The method according to claim 1, characterized in that The method of constructing a celebrity conditional regularization loss function based on the ability of celebrity names to follow text input instructions to optimize the re-encoded character image to obtain a final personalized character image also includes: Based on celebrity names and text input instructions, a series of reference image sets R are generated using the trained multimodal large model; Construct celebrity conditional regularization loss function: in, is a reference frame randomly selected from the reference image set R according to the text input instruction y, M represents The mask of the area outside the face area.

7. The method according to claim 1, characterized in that Get the feature vector c(y) corresponding to the text input instruction y: c(y)=[E(v1,v2,…,v n )] Among them, E represents the pre-trained text encoder and n represents the number of words.

8. The method according to claim 1, characterized in that Inputting a text input instruction into the multimodal large model to output an image x that conforms to the text semantic information; the method further includes: Define the image compression encoder and decoder as ε and The encoder compresses the input image x into the latent space to obtain z = ε(x), and the decoder maps the eigenvalues ​​of the latent space back to the image space to obtain 9. The method according to claim 8, characterized in that By compressing the image x, it can be modeled based on the diffusion model: Where t represents the time label of the diffusion model, and ε is the noise randomly sampled at the current moment.

10. A personalized image generation system based on rapid optimization of diffusion model, characterized in that: include: The model training module is used to obtain the text feature vector corresponding to the text input instruction, and obtain a multimodal large model through diffusion model and latent space diffusion model training; A specific person image encoding module is used to encode a preset person image into a text space using an encoder to obtain a new text feature, and to generate an image containing a specific person by replacing the text feature and the words in the text feature vector; An image encoding module based on text description conditions, used to encode an image containing a specific person according to different text input instructions using the multimodal large model to obtain an encoded person image; A word vector encoding module based on celebrity names, used for re-encoding the encoded character image based on a preset celebrity name to obtain a re-encoded character image; The character image optimization module is used to construct a celebrity-conditional regularized loss function based on the ability of celebrity names to follow text input instructions, so as to optimize the re-encoded character image to obtain the final personalized character image.