Image generation method and device, storage medium and electronic equipment
By employing a multimodal information-based image generation method, utilizing a mask encoder and diffusion module based on global and local expert modules, the problem of inaccurate image generation in existing technologies is solved, achieving higher accuracy and flexibility.
Patent Information
- Application Number
- CN202510887196.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
The existing AI-based image generation models are not accurate enough to meet user needs.
A multimodal information-based image generation method is adopted, which uses the encoder in the trained image generation model to encode user prompts, including the mask encoder of global and local expert modules, and combines it with the diffusion module to generate the target image.
It improves the accuracy and flexibility of image generation, better meets user needs, and generates images that better match user expectations.
Smart Images

Figure CN120997318A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an image generation method and device, a storage medium, and an electronic device. BACKGROUND
[0002] With the rapid development of artificial intelligence generation models, the technology of generating images using artificial intelligence generation models is becoming more mature. For example, a face image is generated using an artificial intelligence generation model, and the generated face image is applied in the fields of digital identity verification, virtual human construction, content creation, etc. However, the images generated based on the existing artificial intelligence generation models are still not accurate and cannot meet user needs.
[0003] Therefore, the present application provides an image generation method. SUMMARY
[0004] The present application provides an image generation method, device, storage medium, and electronic device to at least partially solve the above problems in the prior art.
[0005] The present application provides the following technical solutions:
[0006] The present application provides an image generation method, which comprises:
[0007] Obtaining user prompt information, wherein the user prompt information comprises at least one of text modal information and semantic mask modal information;
[0008] For each modal information in the user prompt information, inputting the modal information into an encoder corresponding to the modal information in a trained image generation model to encode the modal information through the encoder corresponding to the modal information, and obtaining an information encoding result of the modal information, wherein the mask encoder corresponding to the semantic mask modal information comprises at least one global expert module for extracting global features and at least two local expert modules for extracting local features;
[0009] Taking each information encoding result as a generation condition of the image generation model, and injecting the generation condition into a diffusion module in the image generation model to make the diffusion module output a target image according to the injected generation condition.
[0010] The present application provides an image generation device, which comprises:
[0011] A user prompt information obtaining module is configured to obtain user prompt information, wherein the user prompt information comprises at least one of text modal information and semantic mask modal information;
[0012] The encoding module is configured to input each modality information in the user prompt information into an encoder corresponding to the modality information in the trained image generation model, to encode the modality information by the encoder corresponding to the modality information, and obtain an information encoding result of the modality information, wherein the mask encoder corresponding to the semantic mask modality information comprises at least one global expert module for extracting global features and at least two local expert modules for extracting local features.
[0013] The target image generation module is configured to take the information encoding results as generation conditions of the image generation model, and inject the generation conditions into a diffusion module in the image generation model, so that the diffusion module outputs a target image according to the injected generation conditions, wherein the diffusion module is obtained based on a diffusion model.
[0014] The present specification provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the image generation method.
[0015] The present specification provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the image generation method when executing the program.
[0016] The above at least one technical solution adopted by the present specification can achieve the following beneficial effects:
[0017] In the image generation method provided by the present specification, the image generation model is deployed with encoders corresponding to multiple modalities of information, so that the image generation model can receive user prompt information of multiple modalities, and the target image generated based on the multiple modalities of information is more accurate. In addition, for each modality of information, the corresponding encoder is used for encoding, which is more targeted and the obtained encoding result is more accurate, so as to improve the accuracy of the target image obtained based on the encoding result. Furthermore, the local expert modules included in the mask encoder corresponding to the semantic mask modality information can extract corresponding features for the corresponding local positions in the semantic mask modality information, and extract global features of the semantic mask based on the global expert module, which ensures the accuracy of the extracted global features and the accuracy of the extracted local features, and the accuracy of the target image obtained based on the global features and the local features is higher. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings explained herein are used to provide further understanding of the present specification, and form a part of the present specification. The illustrative embodiments of the present specification and their descriptions serve to explain the present specification, and do not constitute an improper limitation on the present specification. In the drawings:
[0019] Figure 1 A flowchart of an image generation method provided for the specification of the present application is shown in the figure;
[0020] Figure 2 An interaction diagram between a server and a terminal of a user provided for the specification of the present application is shown in the figure;
[0021] Figure 3 A diagram of semantic mask modal information provided for the specification of the present application is shown in the figure;
[0022] Figure 4 An internal structure diagram of an image generation model provided for the specification of the present application is shown in the figure;
[0023] Figure 5 A structure diagram of a mask encoder provided for the specification of the present application is shown in the figure;
[0024] Figure 6 A structure diagram of another mask encoder provided for the specification of the present application is shown in the figure;
[0025] Figure 7 A structure diagram of a diffusion module provided for the specification of the present application is shown in the figure;
[0026] Figure 8 A training flowchart of an image generation model provided for the specification of the present application is shown in the figure;
[0027] Figure 9 A data flowchart in a mask encoder provided for the specification of the present application is shown in the figure;
[0028] Figure 10 An image generation device diagram provided for the specification of the present application is shown in the figure;
[0029] Figure 11 An electronic device diagram corresponding to the Figure 1 provided for the specification of the present application is shown in the figure. DETAILED DESCRIPTION
[0030] In order to make the purposes, technical solutions and advantages of the specification of the present application clearer, the technical solutions of the specification of the present application will be described clearly and completely below in combination with specific embodiments of the specification of the present application and corresponding figures. Obviously, the described embodiments are only some of the embodiments of the specification of the present application, but not all the embodiments. Based on the embodiments in the specification of the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0031] In order to better understand and illustrate the solutions provided by the specification of the present application, some technical terms involved in the specification of the present application will be briefly introduced below.
[0032] The semantic mask technology refers to a technology for identifying the object category to which each pixel in an image belongs. By assigning different labels or values to different parts of the image, it helps to distinguish and identify different regions of the image, such as people, backgrounds, objects, etc.
[0033] Diffusion models are a class of generative models that simulate the process of gradually adding noise to data (forward diffusion) and removing noise (inverse generation) to generate high-quality data.
[0034] The execution subject of the present application specification is various computing devices that can execute the image generation method provided by the present application specification, such as a single server, a server cluster, etc. The trained image generation model is deployed in the computing device to provide services for users, and the computing device can communicate with the terminal of the user. For ease of illustration, the server is taken as the execution subject for illustration.
[0035] The technical solutions provided by the embodiments of the present application specification will be described in detail below with reference to the accompanying drawings.
[0036] Figure 1 A flowchart of an image generation method provided by the present application specification is shown, which specifically includes steps S100-S104.
[0037] S100: Obtain user prompt information.
[0038] Figure 2 An interaction diagram between the server and the terminal of the user provided by the present application specification is shown, as shown in Figure 2 .
[0039] The user can send the user prompt information to the server through the terminal, and the server receives the user prompt information sent by the terminal, so as to subsequently generate a target image meeting the user's demand based on the user prompt information. The target image can be a face image, a landscape image, etc., which is not limited by the present application specification.
[0040] The user prompt information is data indicating the user's demand for the target image. The user prompt information can include multiple modalities, specifically at least one of a text modality and a semantic mask modality. In other words, the user prompt information includes at least one of text modality information and semantic mask modality information. The text modality information can include several texts, and the semantic mask modality information can include at least one global mask, which is composed of at least two local masks. It can be understood that the image generation model in the present application specification can generate a target image through user prompt information of multiple modalities, or generate a target image through user prompt information of a single modality, which is flexible and practical.
[0041] Figure 3 An illustrative diagram of semantic mask modal information provided in the specification is shown in FIG. 1. Figure 3
[0042] In the specification, taking a generated target image as a face image as an example, the semantic mask modal information includes at least one face global mask, and the face global mask is composed of at least two face local masks, specifically, the face global mask can be composed of face local masks of parts such as face, hair, nose, etc.
[0043] S102: For each modal information in the user prompt information, input the modal information into an encoder corresponding to the modal information in the trained image generation model to encode the modal information through the encoder corresponding to the modal information, and obtain an information encoding result of the modal information.
[0044] Figure 4 An internal structure diagram of the image generation model provided in the specification is shown in FIG. 1. Figure 4
[0045] The image generation model includes an encoder corresponding to each modal information, specifically including a mask encoder and a text encoder, and the specification does not limit the specific type of the text encoder, which can encode the text.
[0046] Therefore, when inputting the user prompt information into the image generation model, for each modal information in the user prompt information, the modal information can be input into an encoder corresponding to the modal information in the trained image generation model to encode the modal information through the encoder corresponding to the modal information, and obtain an information encoding result of the modal information.
[0047] Specifically, when the modal information in the user prompt information is text modal information, the text modal information is input into the text encoder in the image generation model, and the text encoder encodes the text modal information to obtain an information encoding result corresponding to the text modal information. When the modal information in the user prompt information is semantic mask modal information, the semantic mask modal information is input into the mask encoder in the image generation model, and the mask encoder encodes the semantic mask modal information to obtain an information encoding result corresponding to the semantic mask modal information.
[0048] In the image generation model, a corresponding encoder is deployed for each modality information, so that the image generation model can receive information of multiple modalities, and the accuracy of the target image generated based on the multi-modal information is higher. In addition, for each modality information, a dedicated encoder is used for encoding, and the obtained information encoding result is more accurate. This is because each modality has unique features and structures, and a specially designed encoder can more effectively capture the characteristics of the corresponding modality information.
[0049] Figure 5 The structure diagram of the mask encoder provided in the present application is shown in Figure 5
[0050] The mask encoder corresponding to the semantic mask modality information includes at least one global expert module for extracting global features, at least two local expert modules for extracting local features, an image encoding module, and a dynamic gating network. This structure including global expert modules and local expert modules can also be called a mixture of global and local experts (MoGLE) structure. Among them, one or more multi-layer perceptrons can be included in the global expert module and the local expert module to extract the features of the semantic mask modality information.
[0051] When the semantic mask modality information is input into the mask encoder, the global mask in the semantic mask modality information and all local masks constituting the global mask can be input into the image encoding module respectively to obtain the image encoding result corresponding to each mask; the image encoding result corresponding to the global mask is input into the global expert module to obtain the first feature of the global mask output by the global expert module. And for each local mask, the local mask is input into the local expert module corresponding to the local mask to obtain the second feature of the local mask output by the local expert module; based on the first feature and each second feature, the information encoding result of the semantic mask modality information is obtained through the dynamic gating network.
[0052] Taking the global mask included in the semantic mask modal information as a face global mask, the face global mask is composed of face local masks of parts such as face, hair, nose, etc. When the face semantic mask is input into the mask encoder, the face global mask and a plurality of face local masks can be input into an image encoding module. The image encoding module can include an image encoder, which can be a variational autoencoder (VAE). The image encoder can perform image encoding operations on the face global mask and the plurality of face local masks in parallel to obtain image encoding results corresponding to the face global mask and the plurality of face local masks. Of course, the image encoding operations on the face global mask and the plurality of face local masks can also be performed in series. That is, the face global mask and the plurality of face local masks are input into the image encoder in a predetermined order, so that the image encoder performs image encoding operations on the input masks in the predetermined order.
[0053] Subsequently, the image encoding result corresponding to the face global mask is input into the global expert module. The global expert module captures features of the image encoding result corresponding to the face global mask, focuses on extracting global features of the face global mask, and obtains first features of the face global mask. The image encoding result corresponding to each face local mask is input into the local expert module corresponding to the face local mask. For example, the image encoding result corresponding to the face mask is input into the local expert module corresponding to the face data, and the image encoding result corresponding to the nose mask is input into the local expert module corresponding to the nose data. The local expert module extracts local features and outputs second features. Finally, based on the first features and the second features, the information encoding result of the face semantic mask is obtained through a dynamic gating network.
[0054] By using the global expert module and the plurality of local expert modules to process the image encoding result, each expert module focuses on extracting features of the corresponding mask, and the obtained features are more representative and can better represent the corresponding mask. Therefore, the target image obtained based on the features extracted by the expert module meets the user's demand. In other words, for the global expert module, it focuses on capturing global features of the image encoding result corresponding to the global mask to obtain more global features. For the local expert module, it focuses on extracting local features in the image encoding result corresponding to the local mask to achieve fine feature capture. By extracting features at different levels, more feature information that can represent the user's demand is obtained, and a more accurate target image is obtained.
[0055] That is to say, the local expert module included in the mask encoder corresponding to the semantic mask modal information can extract corresponding features for the corresponding local position in the semantic mask modal information, and extract global features of the semantic mask based on the global expert module, which ensures the accuracy of the extracted global features and the accuracy of the extracted local features, and the accuracy of the target image based on the global features and the local features is higher.
[0056] It can be understood that by mixing the expert structure, the semantic mask modal information is split by region, and a dedicated local expert modeling submodule is introduced, which effectively realizes the decoupling modeling of the structure semantics, and at the same time, the global expert is used to maintain overall consistency and avoid image fragmentation caused by local control, and the modeling freedom and regional accuracy are higher.
[0057] Figure 6 Another structure diagram of a mask encoder provided in the specification of the present application is shown in Figure 6 .
[0058] The image encoding module in the mask encoder can include a plurality of image encoders, and the number of image encoders in the image encoding module is the sum of the number of global masks and local masks. That is to say, one mask corresponds to one image encoder, so when the face semantic mask is input into the mask encoder, the face global mask is input into an image encoder, and for each face local mask, the face local mask is input into the image encoder, so that the image encoder performs image encoding operation. One mask corresponds to one image encoder can improve the image encoding efficiency, and then improve the efficiency of obtaining the target image, reduce the user waiting time, and improve the user experience.
[0059] S104: encode each information as a generation condition of the image generation model, and inject the generation condition into a diffusion module in the image generation model, so that the diffusion module outputs a target image according to the injected generation condition, wherein the diffusion module is obtained based on a diffusion model.
[0060] In the specification of the present application, Figure 7 a structure diagram of a diffusion module provided in the specification of the present application is shown in Figure 7 . The diffusion module includes an adapter, a pre-trained diffusion model and an image decoding module. The adapter is obtained based on a low-rank adaptation (Low-Rank Adaptation Adapter, LoRA) mechanism.
[0061] The server can input the generation condition into the adapter to make the adapter perform fusion operation on each information encoding result corresponding to each modality information in the generation condition to obtain a fused generation condition when generating a target image by taking each information encoding result as a generation condition of an image generation model, injecting the generation condition into a diffusion module in the image generation model, and obtaining a target image output by the diffusion module.
[0062] The condition fusion is needed because the space that can be injected by the diffusion model is limited, and multiple generation conditions need to be injected in the limited space to improve the accuracy of generating a target image. It can be understood that the adapter can perform format conversion, enhancement or injection of specific content information on the input features.
[0063] It should be noted that there are adapters corresponding to different modality information, that is, the number of adapters in the diffusion module can be set as needed. Specifically, in the present application, when the information encoding result input into the adapter only includes information encoding result of text modality information, the information encoding result of the text modality information is input into the adapter corresponding to the text modality information. When the information encoding result input into the adapter only includes information encoding result of semantic mask modality information, the information encoding result of the semantic mask modality information is input into the adapter corresponding to the semantic mask modality information. When the information encoding result input into the adapter includes information encoding result of semantic mask modality information and information encoding result of text modality information, the information encoding result of semantic mask modality information and information encoding result of text modality information can be input into the adapter corresponding to the mixed modality information.
[0064] Then, the fused generation condition is injected into the diffusion model to make the diffusion model gradually perform noise removal operation on the preset standard noise feature to obtain a denoised target image feature. The diffusion model can take the image feature obtained after performing the preset number of noise removal operations as the target image feature.
[0065] Finally, the server inputs the denoised target image feature into the image decoder to obtain a target image output by the image decoder. The preset standard noise feature is obtained by encoding a standard noise image. Generally, the diffusion model gradually denoises the standard noise image to obtain a denoised image. In the present application, the standard noise feature is in a latent space dimension, the standard noise image is in a pixel dimension, and the dimension of the standard noise feature is lower than that of the standard noise image. Therefore, the diffusion model needs to process less data when denoising the standard noise feature, thereby reducing the computational cost and time, and being more efficient.
[0066] The standard noise feature is removed by a diffusion model to restore the image feature closest to the user's demand, so as to obtain a more accurate target image. In addition, the diffusion model simulates a process of gradually adding noise, that is, a forward diffusion process, and then learns how to reverse this process, that is, a backward generation process. This step-by-step noise removal method enables the model to finely adjust the value of each pixel, thereby generating a data sample with rich details and high quality. Therefore, the diffusion model can more accurately control each step of the generation process to ensure that the final output image is more in line with the user's demand. In addition, using the diffusion model for image generation has higher stability, because the diffusion model is designed based on Markov chain, and each small step of noise removal is a relatively simple task, so the entire generation process is relatively stable.
[0067] Based on Figure 1 According to the image generation method shown in the image generation model, multiple modal information corresponding encoders are deployed, so that the image generation model can receive multiple modal user prompt information, and the target image generated based on the multiple modal information has higher accuracy. In addition, for each modal information, the corresponding encoder is used for encoding, which is more targeted and the encoding result is more accurate, thereby improving the accuracy of the target image obtained based on the encoding result. In addition, the local expert module included in the mask encoder corresponding to the semantic mask modal information can extract the corresponding features for the corresponding local position in the semantic mask modal information, and extract the global features of the semantic mask based on the global expert module, which ensures the accuracy of the extracted global features and the accuracy of the extracted local features. Therefore, the target image obtained based on the global features and the local features has higher accuracy.
[0068] For step S102, when the information encoding result of the semantic mask modal information is obtained based on the first feature and the second features through the dynamic gating network, the image encoding result corresponding to the global mask, the first feature and the second features can be input into the dynamic gating network to obtain the weight corresponding to each feature output by the dynamic gating network; and the first feature and the second features are weighted and averaged based on the weight of each feature to obtain the information encoding result of the semantic mask modal information.
[0069] The outputs of all experts are fused through a dynamic gating network, which has timing perception ability and has learned how to dynamically adjust the fusion weights of each expert according to the noise amount and spatial position through training, thereby realizing the organic unification of local details and global structure. Therefore, through the dynamic gating network, the weight corresponding to the feature output by each expert module can be predicted, and then each feature is weighted and averaged according to the weight corresponding to each feature to obtain the information encoding result of the semantic mask modal information, thereby improving the accuracy of the target image obtained based on the information encoding result.
[0070] The present application also provides a training method of an image generation model. The execution subject of the training method can be a computing device capable of training the model, such as a server. For the sake of illustration, the server is taken as the execution subject of the training of the image generation model.
[0071] Figure 8 The training flowchart of the image generation model provided by the present application is shown in FIG. 1. Figure 8
[0072] The snowflake-shaped parameters represent frozen parameters, and the spark-shaped parameters represent learnable parameters. When adjusting the parameters, the frozen parameters do not need to be adjusted, and the learnable parameters are adjusted to improve the training efficiency.
[0073] When training the image generation model, the server can first obtain training sample data, which includes a sample image, a sample semantic mask corresponding to the sample image, and a sample text. Among them, Figure 8 The human face image in the present application is the sample image, Figure 8 The text prompt in the present application is the sample text, Figure 8 The semantic mask in the present application is the sample semantic mask.
[0074] Then, the server can input the sample information of each modality in the training sample data into the corresponding encoder. Specifically, the sample image is input into the image encoder to obtain the first encoding result Z of the sample image output by the image encoder. The sample text is input into the text encoder in the image generation model to obtain the second encoding result C output by the text encoder. p The sample semantic mask is input into the mask encoder in the image generation model to obtain the third encoding result C output by the mask encoder. m
[0075] It should be noted that the image encoder in the image encoding module in the image encoder and the mask encoder can be the same encoder, or can be different encoders, which is not limited in the present application.
[0076] The training of the image generation model includes a forward process of adding Gaussian noise to the latent image token and a denoising process of learning to restore the real image. Therefore, the server can also add noise to the first encoding result according to a preset time step to obtain the first encoding result Z after noise addition. t The preset time step t can be randomly obtained, that is, the preset time step can be dynamically changed, and the noise amount for adding noise to the first encoding result is dynamically changed each time. By adding different noise amounts, the image generation model has stronger noise prediction capability, so that when applied, the noise amount in the user prompt information can be accurately predicted and denoising operation is performed, and a more accurate target image is obtained.
[0077] The first encoding result, the second encoding result and the third encoding result after adding noise are injected into a diffusion module included in the image generation model to obtain a predicted noise amount ∈. Specifically, the first encoding result, the second encoding result and the third encoding result after adding noise are input into an adapter, so that the adapter performs fusion operation on the first encoding result, the second encoding result and the third encoding result after adding noise to obtain a fused sample encoding result. The fused sample encoding result is injected into a diffusion model to make the diffusion model output the predicted noise amount. By taking the information encoding result corresponding to each modality information output by each encoder and the information encoding result of each region in the semantic mask as a generation condition, the diffusion model is guided to generate an image conforming to the sample semantic mask and the sample text.
[0078] The image generation model is trained according to the predicted noise amount and the noise amount for adding noise to the first encoding result. Specifically, according to the predicted noise amount and the noise amount for adding noise to the first encoding result, a loss is determined, and the image generation model is trained with the loss reduction as the training target. The loss function used when determining the loss can be a mean square error loss function, of course, other loss functions can also be used, which are not limited in the present application. In the present application, the diffusion model needs to learn how to control the diffusion model to denoise the standard noise features through the adapter to obtain an image conforming to the sample semantic mask and the sample text as much as possible, so the model parameters of the image generation model need to be adjusted with the loss reduction as the training target.
[0079] When training the image generation model, the adjustable parameters include the parameters of the mask encoder and the parameters of the adapter, wherein the parameters of the mask encoder can include the parameters of each expert module and the parameters of the dynamic gating network.
[0080] The efficient fine-tuning of the model is realized by adding an adapter based on a low-rank matrix to the pre-trained diffusion model, which significantly reduces the number of parameters that need to be adjusted, reduces the computational cost and resource demand, while maintaining the performance and effectiveness of the model. Therefore, when training the model, the pre-trained diffusion model does not need to be adjusted, and only the low-rank matrix needs to be adjusted to complete the training. In other words, the adapter not only can format conversion, enhance or inject specific content information to the input features, but also can realize the effect of reducing the number of modified parameters by modifying the parameters of the adapter during training without modifying the parameters of the diffusion model.
[0081] In addition, in the training stage of the present application, there is no need to generate images in the pixel space, which significantly improves the computational efficiency and generation quality.
[0082] It should be noted that in order to improve the generalization ability of the image generation model, the image generation model can not only generate target images through single-modal user prompt information, but also generate target images through multi-modal user prompt information. Before inputting the sample information of each modality in the training sample data into the corresponding encoder, the server can also perform a zero operation on the sample information of each modality through a random probability. That is to say, each modality of sample information has a probability of being set to a default value, and the default value can be 0. The foregoing process can be referred to as a conditional random dropout strategy. When only one modality of sample information is not set to a default value, the feature received by the image generation model is a single-modal sample feature. When at least two modalities of sample information are not set to a default value, the feature received by the image generation model is a multi-modal sample feature.
[0083] In other words, the image generation model uses a conditional random dropout strategy during training to enhance the adaptability to the control signal, thereby supporting the free combination or omission of any conditional input in the application stage, greatly improving the practicability and flexibility of the image generation model, and laying a foundation for future human-computer interactive generation systems.
[0084] Figure 9 The data flow chart in the mask encoder provided in the present application is shown in Figure 9 .
[0085] When inputting the sample semantic mask into the mask encoder in the image generation model to obtain the third encoding result output by the mask encoder, the precise control of each region is realized by assigning corresponding expert modules to different regions of one mask.
[0086] Specifically, the server can input the sample global mask and the sample local mask in the sample semantic mask into an image coding module in the mask encoder respectively to obtain a sample coding result corresponding to each sample mask. The image coding module can include one or more image encoders, Figure 9 The weight sharing in the image coding module refers to that the parameters in each image encoder can be the same. The sample coding result belongs to an L*d matrix space in a real number domain, the sample coding result corresponding to the sample global mask is The sample coding result corresponding to the sample local mask includes n is the number of sample local masks.
[0087] The sample coding result corresponding to the sample global mask is input into a global expert module in the mask encoder to obtain a first sample feature of the sample global mask output by the global expert module; and for each sample local mask, the sample coding result corresponding to the sample local mask is input into a local expert module corresponding to the sample local mask in the mask encoder to obtain a second sample feature of the sample local mask output by the local expert module.
[0088] Then, the sample coding result corresponding to the sample global mask, the first sample feature, each second sample feature, a preset time step and the first coding result after noise addition are input into a dynamic gating network to obtain a weight corresponding to each sample feature output by the dynamic gating network; the first sample feature and each second sample feature are weighted and averaged based on the weight of each sample feature to obtain a third coding result of the sample semantic mask. Figure 9 The global mask token in the image coding module refers to the sample coding result corresponding to the sample global mask, and the image token after noise addition refers to the first coding result after noise addition.
[0089] Through the dynamic gating network, it is intelligently selected which features output by the experts are fused according to the current diffusion time step and spatial position. Compared with the static control mode, this mechanism has timing perception and spatial adaptability, can dynamically focus on the key area, and realizes optimization control in different stages in the whole diffusion process. The spatial position refers to that each pixel point is allocated a corresponding weight, the weight changes with the change of the time step, and through training, the parameters of the expert module are adjusted so that the expert module learns the best corresponding relationship between the time step and the weight.
[0090] It can be understood that, in the present application, the training process is carried out in the latent space without generating images, that is, without performing in the pixel space, and the image encoder is combined to realize efficient modeling and high-quality image restoration, and to improve the generation efficiency and stability. The image generation method provided in the present application introduces a hybrid expert structure, combines the time sequence characteristics in the diffusion generation process, constructs a collaborative network architecture composed of global experts and local experts, and uses a dynamic gating network to realize dynamic fusion and selection of different experts at different times and spatial positions, thereby significantly improving the model's ability in semantic understanding, local control, structural consistency and multi-modal adaptability. The image generation method provided in the present application can not only be applied to the task of generating controllable high-quality and diversified human face images, but also provides training data with more control and authenticity for generative large models, which helps to build a safe and reliable face recognition and synthesis system, and meets the core demand of intelligent generation capability for future multi-condition interactive man-machine systems.
[0091] For the field of human face image generation, the image generation method provided in the present application introduces a dynamic collaborative expert module and a diffusion control mechanism in the generation process, realizes fine-grained regulation of multiple key semantic regions in human face images such as hairstyle, glasses, facial contour, etc. The method supports interactive joint control of multi-modal input conditions such as text and semantic masks, and introduces a dynamic gating network to realize dynamic evolution of expert selection with time and space, thereby greatly enhancing the controllable ability of image structure and semantics while maintaining high fidelity. The method supports multi-condition interactive control, accurately controls specific regions and attributes of human face images, and can stably generate high-quality images under multi-modal input, with high adaptability. The global expert guarantees the overall structural consistency, and the local expert improves the quality of regional details. It supports step-by-step editing and real-time feedback of control conditions, and improves the interactive experience of the generation process. The method can be applied to multiple downstream tasks such as data augmentation, face forgery detection and image editing, and has strong universality. Of course, the effects achieved by the present application can be realized not only in generating human face images, but also in generating other images.
[0092] The above is the image generation method provided by one or more embodiments of the present application. Based on the same idea, the present application also provides a corresponding image generation device, as shown in Figure 10 The device comprises:
[0093] The user prompt information acquisition module 1000 is configured to acquire user prompt information, wherein the user prompt information comprises at least one of text modal information and semantic mask modal information;
[0094] The encoding module 1002 is configured to input each modality information in the user prompt information into an encoder corresponding to the modality information in a trained image generation model, to encode the modality information by the encoder corresponding to the modality information, and to obtain an information encoding result of the modality information, wherein the mask encoder corresponding to the semantic mask modality information comprises at least one global expert module for extracting global features and at least two local expert modules for extracting local features.
[0095] The target image generation module 1004 is configured to take the information encoding results as generation conditions of the image generation model, to inject the generation conditions into a diffusion module in the image generation model, to make the diffusion module output a target image according to the injected generation conditions, and to obtain the target image.
[0096] Optionally, the semantic mask modality information comprises at least one global mask, the global mask is composed of at least two local masks, and the mask encoder corresponding to the semantic mask modality information further comprises an image encoding module and a dynamic gating network.
[0097] When the modality information is semantic mask modality information, the encoding module 1002 is specifically configured to input a global mask in the semantic mask modality information and all local masks constituting the global mask into the image encoding module respectively, to obtain image encoding results corresponding to the masks.
[0098] The image encoding result corresponding to the global mask is input into the global expert module to obtain first features of the global mask output by the global expert module, and each local mask is input into a local expert module corresponding to the local mask to obtain second features of the local mask output by the local expert module.
[0099] Based on the first features and the second features, the dynamic gating network is used to obtain an information encoding result of the semantic mask modality information.
[0100] Optionally, the encoding module 1002 is specifically configured to input the image encoding result corresponding to the global mask, the first features and the second features into the dynamic gating network to obtain weights corresponding to each feature output by the dynamic gating network.
[0101] Based on the weights of each feature, the first features and the second features are weighted and averaged to obtain the information encoding result of the semantic mask modality information.
[0102] Optionally, the diffusion module comprises an adapter, a diffusion model and an image decoder.
[0103] The target image generation module 1004 is specifically configured to input the generation condition into the adapter, so that the adapter performs fusion operation on each information encoding result corresponding to each modality information in the generation condition to obtain a fused generation condition.
[0104] The fused generation condition is injected into the diffusion model, so that the diffusion model performs noise removal operation on the preset standard noise feature step by step to obtain a denoised target image feature.
[0105] The denoised target image feature is input into the image decoder to obtain a target image output by the image decoder.
[0106] Optionally, the device further comprises:
[0107] The training module 1006 is configured to obtain training sample data, wherein the training sample data comprises a sample image, a sample semantic mask corresponding to the sample image, and a sample text.
[0108] The sample image is input into an image encoder to obtain a first encoding result of the sample image output by the image encoder; and the first encoding result is added with noise according to a preset time step to obtain a first encoding result added with noise.
[0109] The sample text is input into a text encoder in the image generation model to obtain a second encoding result output by the text encoder, and the sample semantic mask is input into a mask encoder in the image generation model to obtain a third encoding result output by the mask encoder.
[0110] The first encoding result added with noise, the second encoding result, and the third encoding result are injected into a diffusion module included in the image generation model to obtain a predicted noise quantity output by the diffusion module.
[0111] The image generation model is trained according to the predicted noise quantity and a noise quantity with which the first encoding result is added with noise.
[0112] Optionally, the training module 1006 is specifically configured to perform zero operation on sample information of each modality by random probability before the sample information of each modality in the training sample data is input into a corresponding encoder.
[0113] Optionally, the training module 1006 is specifically configured to input a sample global mask and a sample local mask in the sample semantic mask into an image encoding module in the mask encoder respectively to obtain a sample encoding result corresponding to each sample mask.
[0114] input the sample encoding result corresponding to the sample global mask into a global expert module in the mask encoder, to obtain first sample features of the sample global mask output by the global expert module; and for each sample local mask, input the sample encoding result corresponding to the sample local mask into a local expert module corresponding to the sample local mask in the mask encoder, to obtain second sample features of the sample local mask output by the local expert module;
[0115] Based on the first sample features and each second sample feature, a third encoding result of the sample semantic mask is obtained through a dynamic gating network in the mask encoder.
[0116] Optionally, the training module 1006 is specifically configured to input the sample encoding result corresponding to the sample global mask, the first sample features, each second sample feature, the preset time step, and the first encoding result after adding noise into the dynamic gating network, to obtain a weight corresponding to each sample feature output by the dynamic gating network.
[0117] Based on the weight of each sample feature, the first sample features and each second sample feature are weighted and averaged, to obtain the third encoding result of the sample semantic mask.
[0118] The present specification also provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the image generation method described above. Figure 1 The image generation method is provided.
[0119] The present specification also provides an electronic device. Figure 11 As shown in the structural schematic diagram of the electronic device, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Figure 11 The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs, to implement the image generation method described above. Figure 1 Of course, in addition to the software implementation manner, the present specification does not exclude other implementation manners, such as a logic device or a combination of software and hardware, that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.
[0120] In the 1990s, it was relatively easy to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has evolved, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flows into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming the PLD, rather than by ordering a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in the above-mentioned hardware description languages and programmed into an integrated circuit, a hardware circuit implementing the logical method flow can be easily obtained.
[0121] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to being implemented in pure computer readable program code, the controller can equally well be implemented to perform the same functions using logic gates, switches, an application specific integrated circuit, a programmable logic controller and an embedded microcontroller, etc. by means of a logical programming of the method steps. The controller can thus be considered as a hardware component, and the means comprised therein for performing the various functions can be considered as structures within the hardware component. Alternatively, the means for performing the various functions can even be considered as both a software module implementing the method and a structure within the hardware component.
[0122] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0123] For the sake of description, the above apparatuses are described in various units by functions respectively. Of course, the functions of the units can be implemented in one or more software and / or hardware in the implementation of the present application.
[0124] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer usable program code.
[0125] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0126] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0127] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0128] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0129] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory. The memory is an example of computer-readable media.
[0130] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0131] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device that includes the element.
[0132] Those skilled in the art will appreciate that embodiments of the present application specification can be provided as methods, systems or computer program products. Therefore, the present application specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The present application specification can be described in the general context of computer-executable instructions, such as program modules, executed by computers. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0134] Each of the embodiments in the specification of the present application is described in a progressive manner, and the same or similar parts between the embodiments can be mutually referred to, and each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the description of the method embodiments.
[0135] The above only describes the embodiments of the specification of the present application, and is not used to limit the specification of the present application. The specification of the present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the specification of the present application shall be included in the scope of claims of the present application.
Claims
1. An image generation method, the method comprising: obtaining user prompt information, the user prompt information comprising at least one of text modal information and semantic mask modal information; for each modal information in the user prompt information, inputting the modal information into an encoder corresponding to the modal information in a trained image generation model to encode the modal information through the encoder corresponding to the modal information, to obtain an information encoding result of the modal information, wherein a mask encoder corresponding to the semantic mask modal information comprises at least one global expert module for extracting global features and at least two local expert modules for extracting local features; taking each information encoding result as a generation condition of the image generation model, and injecting the generation condition into a diffusion module in the image generation model, so that the diffusion module outputs a target image according to the injected generation condition, wherein the diffusion module is obtained based on a diffusion model.
2. The method of claim 1, wherein the semantic mask modal information comprises at least one global mask, the global mask being composed of at least two local masks, and the mask encoder corresponding to the semantic mask modal information further comprises an image encoding module and a dynamic gating network; when the modal information is semantic mask modal information, inputting the modal information into an encoder corresponding to the modal information in a trained image generation model to encode the modal information through the encoder corresponding to the modal information, to obtain an information encoding result of the modal information, specifically comprising: inputting the global mask in the semantic mask modal information and all local masks constituting the global mask into the image encoding module respectively to obtain image encoding results corresponding to each mask; inputting the image encoding result corresponding to the global mask into the global expert module to obtain first features of the global mask output by the global expert module; and for each local mask, inputting the local mask into a local expert module corresponding to the local mask to obtain second features of the local mask output by the local expert module; based on the first features and the second features, obtaining the information encoding result of the semantic mask modal information through the dynamic gating network.
3. The method of claim 2, wherein based on the first features and the second features, the information encoding result of the semantic mask modal information is obtained through the dynamic gating network, specifically comprising: inputting the image encoding result corresponding to the global mask, the first features and the second features into the dynamic gating network to obtain weights corresponding to each feature output by the dynamic gating network; based on the weights of each feature, performing weighted average on the first features and the second features to obtain the information encoding result of the semantic mask modal information.
4. The method of claim 1, wherein the diffusion module comprises an adapter, a diffusion model and an image decoder; injecting the generation condition into the diffusion module in the image generation model so that the diffusion module outputs a target image according to the injected generation condition, specifically comprising: inputting the generated condition into the adapter, so that the adapter performs fusion operation on each information coding result corresponding to each modality information in the generated condition to obtain a fused generated condition; injecting the fused generated condition into the diffusion model, so that the diffusion model performs noise removal operation on the preset standard noise feature step by step to obtain a denoised target image feature; inputting the denoised target image feature into the image decoder to obtain a target image output by the image decoder.
5. The method of claim 1, wherein the image generation model is trained, specifically comprising: obtaining training sample data, the training sample data comprising a sample image, a sample semantic mask corresponding to the sample image, and a sample text; inputting the sample image into an image encoder to obtain a first encoding result of the sample image output by the image encoder; and adding noise to the first encoding result according to a preset time step to obtain a first encoding result after adding noise; inputting the sample text into a text encoder in the image generation model to obtain a second encoding result output by the text encoder, and inputting the sample semantic mask into a mask encoder in the image generation model to obtain a third encoding result output by the mask encoder; injecting the first encoding result after adding noise, the second encoding result, and the third encoding result into a diffusion module included in the image generation model to obtain a predicted noise quantity output by the diffusion module; training the image generation model according to the predicted noise quantity and a noise quantity used to add noise to the first encoding result.
6. The method of claim 5, before inputting sample information of each modality in the training sample data into a corresponding encoder, the method further comprises: performing zero operation on the sample information of each modality through random probability.
7. The method of claim 5, inputting the sample semantic mask into the mask encoder in the image generation model to obtain a third encoding result output by the mask encoder, specifically comprising: inputting a sample global mask and a sample local mask in the sample semantic mask into an image encoding module in the mask encoder respectively to obtain a sample encoding result corresponding to each sample mask; inputting a sample encoding result corresponding to the sample global mask into a global expert module in the mask encoder to obtain a first sample feature of the sample global mask output by the global expert module; and for each sample local mask, inputting a sample encoding result corresponding to the sample local mask into a local expert module corresponding to the sample local mask in the mask encoder to obtain a second sample feature of the sample local mask output by the local expert module; based on the first sample feature and each second sample feature, obtaining a third encoding result of the sample semantic mask through a dynamic gating network in the mask encoder.
8. The method of claim 7, based on the first sample feature and each second sample feature, obtaining a third encoding result of the sample semantic mask through a dynamic gating network in the mask encoder, specifically comprising: input the sample encoding result corresponding to the sample global mask, the first sample feature, each second sample feature, the preset time step and the first encoding result after adding noise into the dynamic gating network to obtain a weight corresponding to each sample feature output by the dynamic gating network; perform weighted average on the first sample feature and each second sample feature based on the weight of each sample feature to obtain a third encoding result of the sample semantic mask. 9.A computer readable storage medium, the storage medium storing a computer program, the computer program being executed by a processor to implement the method in any one of claims 1 to 8. 10.An electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the processor implementing the method in any one of claims 1 to 8 when executing the program.