Image generation method, apparatus, device, storage medium, and program product

By introducing sample data and region mask reinforcement learning into the text image model, the problem of face differentiation generation in multi-person images is solved, and high-quality individual differentiation image generation is achieved.

CN122115622APending Publication Date: 2026-05-29CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing text-based image models struggle to capture individual differences when generating images of multiple people, resulting in low usability of the generated images.

Method used

By constructing sample data containing sample text, sample image pairs, and face region masks, a reinforcement learning framework is used to train the pre-trained diffusion model. A fusion loss function of global and local perceptual loss is introduced to optimize facial feature differentiation.

Benefits of technology

It improves the individual differences of faces and the overall image quality in multi-person images, thereby enhancing the usability of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115622A_ABST
    Figure CN122115622A_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field and provides an image generation method, device and equipment, a storage medium and a program product. The method comprises the following steps: acquiring input text; inputting the input text into an image generation model to obtain an image output by the image generation model; wherein the image generation model is obtained by performing reinforcement training on a pre-trained diffusion model based on sample data; any sample data comprises at least one of sample text, a sample image pair generated based on the sample text, and a mask of a human face region in the sample image pair; the sample image pair comprises a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; and a loss function of the reinforcement training is fused with a global loss of the image and a local perception loss based on the mask of the human face region in the image. The image generation model can output an image with differentiated human faces and overall image quality based on the input text, thereby improving the usability of the image generated based on the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image generation method, apparatus, device, storage medium, and program product. Background Technology

[0002] Current mainstream text-based graph models are typically based on a conditional diffusion model structure, mainly consisting of three core modules: a text encoder that encodes the input natural language prompt into a semantic vector, serving as conditional information for the diffusion model; a variational autoencoder (VAE) that performs image encoding and decoding operations, moving the generation process to the latent space and significantly reducing computational cost; and a denoising network, which is the core of the text-based graph model, typically based on a U-Net or transformer architecture, which, guided by the text encoding, is responsible for gradually removing noise and restoring the image representation in the latent space.

[0003] However, in practical applications, it has been found that when existing text-based image models generate images of multiple people, the facial features of different subjects are often highly similar or even completely identical, making it difficult to reflect individual differences.

[0004] Furthermore, although some progress has been made in addressing the diversity issue of text-based image models, and these processing methods each have their advantages in improving the diversity of generated images, they still have significant limitations in the differential generation of faces. As a result, the usability of the generated images is low. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the related art. To this end, this application proposes an image generation method, apparatus, device, storage medium, and program product to solve the problem that traditional text-based image models have significant limitations in the differential generation of human faces, thereby improving the usability of the generated images.

[0006] The image generation method according to the first aspect of this application includes: Get the input text; The input text is fed into the image generation model to obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0007] According to one embodiment of this application, the loss function for reinforcement training is obtained based on the following method: The loss function of the Diffusion-DPO algorithm is used as the global loss of the image and fused with the local perceptual loss based on the face region mask in the image to obtain the loss function of the reinforcement training.

[0008] According to one embodiment of this application, the local perceptual loss is obtained by introducing a face region mask in the image based on the global loss.

[0009] According to one embodiment of this application, the step of fusing the loss function of the Diffusion-DPO algorithm as the global loss of the image with a local perceptual loss based on a face region mask in the image includes: Determine the balance coefficient; Based on the balance coefficient, the global loss and the local perception loss are fused.

[0010] According to one embodiment of this application, the balance coefficient is determined based on the proportion of the face region in the image.

[0011] According to one embodiment of this application, the positive sample image is obtained by performing face swapping on the negative sample image.

[0012] An image generation apparatus according to a second aspect embodiment of this application includes: The acquisition module is used to acquire the input text; The generation module is used to input the input text into the image generation model and obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0013] An electronic device according to a third aspect of this application includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the image generation methods described above.

[0014] According to a fourth aspect of this application, the storage medium is a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the image generation method as described above.

[0015] A computer program product according to a fifth aspect of this application includes a computer program that, when executed by a processor, implements the image generation method as described above.

[0016] The above-described one or more technical solutions in the embodiments of this application have at least the following technical effects: By constructing sample data containing sample text, sample image pairs generated from the sample text, and masks of facial regions within the sample image pairs, and because the sample image pairs include negative sample images generated from the sample text and positive sample images generated from the negative sample images, and provide masks of facial regions within the sample image pairs, differential preference signals and region localization information can be provided simultaneously at the sample level. This allows the model to focus on differential learning of facial features during reinforcement training. Furthermore, when reinforcing the pre-trained diffusion model based on the sample data, facial region masks are introduced to calculate local perceptual loss, thereby improving facial differentiation in multi-person images. The local perceptual loss and global loss are fused as the loss function for reinforcement training, enabling the model to balance local differences and overall consistency. Consequently, the trained image generation model can balance differentiated faces and overall image quality during image generation. Therefore, after obtaining the input text, inputting the input text into the image generation model yields an image that balances differentiated faces and overall image quality, improving the usability of text-based image generation.

[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the image generation method provided in the embodiments of this application.

[0020] Figure 2 This is a schematic diagram of the sample data construction process in the image generation method provided in this application embodiment.

[0021] Figure 3 This is a schematic diagram of the overall technical solution of the image generation method provided in the embodiments of this application.

[0022] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Understandably, existing methods for improving facial diversity include cue word processing and sampling strategies.

[0025] Prompt processing primarily enhances the diversity of generated results by designing varied text descriptions or introducing randomization elements, allowing the model to explore a richer semantic space during generation. These methods typically rely on manual design or automated generation of diverse prompts, offering simplicity and ease of implementation but limiting the potential for increased diversity.

[0026] Sampling strategies involve adjusting the sampling process during the model inference stage. Examples include introducing Conditional-Annealed Diffusion Sampler (CADS) and Classifier-Free Guidance (CFG) parameter adjustments to balance generation quality and diversity, resulting in richer output distributions. These methods do not require modification of model weights, offering advantages such as low cost and ease of implementation, making them an important means of improving the diversity of text-generated images.

[0027] While each of the aforementioned methods has its advantages in enhancing the diversity of generated images, they still have significant limitations in the differentiated generation of faces. The core problem lies in the fact that current methods typically involve global adjustments and generally lack the ability to perceive specific regions of the face. This makes it difficult to achieve diverse control over specific areas of the face, resulting in images that are generally more diverse in appearance but lack sufficient differentiation in local facial features, making it difficult to achieve natural and realistic individual differences.

[0028] Furthermore, the two types of methods that do not require training, namely prompt word processing and sampling strategies, still face problems such as insufficient diversity and impaired generation quality, and cannot fundamentally solve the diversity problem.

[0029] This application proposes an image generation method, apparatus, device, storage medium, and program product, aiming to improve the performance of generative artificial intelligence models in multi-person image generation tasks to address individual facial differences.

[0030] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and regulations of the locality and with authorization from the owner of the relevant device.

[0031] Figure 1 This is a flowchart illustrating the image generation method provided in an embodiment of this application, as shown below. Figure 1 As shown, the image generation method includes: Step 110: Obtain the input text.

[0032] Step 120: Input the input text into the image generation model to obtain the image output by the image generation model.

[0033] The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data. Each sample data includes sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair. The sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image. The loss function of the reinforcement training combines the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0034] It should be noted that the execution subject of the image generation method provided in this application embodiment can be a server or computer device, etc. The computer device can be, for example, a mobile phone, tablet computer, laptop computer, handheld computer, vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.

[0035] It should be noted that this application can pre-build batches of sample data.

[0036] Figure 2 This is a schematic diagram illustrating the process of constructing sample data in the image generation method provided in this application embodiment, such as... Figure 2 As shown, specifically, large batches of prompts containing multiple subjects can be created as sample texts using methods such as knowledge graphs.

[0037] Furthermore, for each sample text, an initial multi-person image can be generated using the text-based image model to be trained, ensuring that the facial regions of each subject in the image are clearly distinguishable. The text-based image model to be trained can be used to generate images containing multiple people based on the input text content; this application does not limit the specific information of the text-based image model.

[0038] Since the raw image model to be trained lacks an explicit mechanism for distinguishing individual identity features, there are often cases where the facial features of different people are highly similar or even nearly identical, making it difficult to reflect the individual differences between people. Therefore, this application can determine each initial multi-person image as a negative sample image (i.e., face homogenization sample).

[0039] To construct training samples with explicit difference signals, this application can further employ face-swapping technology based on face detection and reconstruction to replace or perturb some facial regions in the initial multi-person image (i.e., negative sample image) to generate face version images with obvious differences and determine them as positive sample images (i.e., face difference samples).

[0040] Furthermore, negative sample images and differentiated positive sample images are constructed in a one-to-one correspondence relationship (in this application, sample image pairs can also be referred to as sample pairs), where the former is used as the inferior sample and the latter as the superior sample, and they are used together for preference modeling in the reinforcement learning stage.

[0041] Meanwhile, to achieve subsequent reinforcement learning optimization for region perception, this application further extracts a mask of the face region in each image (including each positive and negative sample image) during the sample pair construction process. This mask can be automatically generated through face detection, key point localization, or semantic segmentation models, and is used to accurately identify the range of the face region of each subject.

[0042] It should be noted that during the reinforcement learning training phase, this mask will be used as the input for region weights to highlight the importance of facial regions in the loss calculation, thereby achieving differentiated guidance for key facial regions.

[0043] Furthermore, this application can use sample text, sample image pairs generated based on sample text, and information such as the mask of the face region in the sample image pair as sample data, thereby providing both "differential preference signals" and "regional localization information" at the sample level. This provides a foundation for subsequent region-aware direct preference optimization (DPO) optimization, enabling the model to focus on the differential learning of facial features during reinforcement learning.

[0044] Furthermore, this application can use sample data to perform preference-driven retraining of the text-to-image model (specifically, a diffusion model that can generate images based on input text after pre-training) based on a reinforcement learning (RL) framework. After training is completed, the image generation model of this application can be obtained.

[0045] It should be noted that this application differs from traditional supervised fine-tuning (SFT) in its retraining (i.e., reinforcement training) based on reinforcement learning.

[0046] SFT optimizes by minimizing the reconstruction error between the model output and the reference image. It is a static, point-to-point supervised learning method that can only improve the overall image clarity or image-text matching, but is difficult to reflect high-level semantic preference goals such as facial differences.

[0047] In contrast, reinforcement learning does not rely on explicit pixel-level labels or fixed targets. Instead, it optimizes the model's generation strategy based on reward or preference signals, thereby directly optimizing the model's behavioral output at a higher semantic level. By introducing reinforcement learning mechanisms, the model no longer simply learns to "reproduce" training data, but can adjust the generation distribution based on feedback signals during the generation process, making it more consistent with the differentiated target features of people in multi-person images.

[0048] Specifically, the loss function for reinforcement training in this application includes the global loss of the image and the local perceptual loss based on the face region mask in the image, thereby forming a reinforcement learning strategy that adopts region perception. This allows the model to perform preference optimization on the difference in noise prediction error between good and bad samples only within the region corresponding to the mask, thereby achieving differentiated guidance for key regions, promoting the model to learn the feature differences of face regions more accurately, and realizing region-level reinforcement adjustment.

[0049] Furthermore, the loss function used in this application for enhanced training is obtained by fusing the global loss of the image and the local perceptual loss based on the face region mask in the image. This allows the model to balance local differences and overall consistency, achieving a balance between global image consistency and local face differentiation. This results in high-fidelity and significantly different multi-person image generation while effectively avoiding model collapse during continuous optimization.

[0050] Furthermore, the trained image generation model is deployed to the execution entity of this application (e.g., a server that can provide text-generated image functionality) so that users can obtain the desired image by inputting text.

[0051] Therefore, users can input specified text into the execution entity of this application to instruct image generation. For example, users can input "generate an image of three young women standing together facing the camera" or "generate an image of several men in suits standing together chatting and laughing," etc.

[0052] Furthermore, the implementing entity of this application can obtain the user's input text and input the input text into the image generation model deployed therein. The image generation model generates an image based on the input text, and finally obtains the image output by the image generation model.

[0053] Furthermore, the entity implementing this application can display the images provided by the image generation model to the user.

[0054] Therefore, compared with traditional enhancement methods, this application adopts a training-side strategy, introducing a region awareness mechanism into the reinforcement learning optimization strategy, and providing differentiated guidance for key face regions in the generation process, enabling the model to learn face regions in a targeted manner, which can effectively improve the individual differences and realism of faces in multi-person images.

[0055] Unlike methods that overlay face-swapping modules after generation, this application fundamentally solves the problem of face homogenization by introducing region-aware reinforcement learning during the training phase of the raw image base model. Post-processing methods can only perform local replacements, which can easily lead to defects such as unnatural edges and inconsistent lighting. In contrast, this application enables the model to distinguish the facial features of different people during the generation stage, thereby improving the facial differences and overall harmony of multi-person images from the source.

[0056] According to the image generation method of this application, sample data is constructed by constructing sample text, sample image pairs generated based on the sample text, and masks of face regions in the sample image pairs. Since the sample image pairs contain negative sample images generated based on the sample text and positive sample images generated based on the negative sample images, and provide masks of face regions in the sample image pairs, differential preference signals and region localization information can be provided simultaneously at the sample level. This allows the model to focus on differential learning of facial features during reinforcement training. Furthermore, when reinforcing training the pre-trained diffusion model based on the sample data, face region masks are introduced to calculate local perceptual loss to improve facial differentiation in multi-person images. The local perceptual loss and global loss are fused as the loss function for reinforcement training, enabling the model to balance local differences and overall consistency. Thus, the trained image generation model can balance differentiated faces and overall image quality when generating images. Therefore, after obtaining the input text, inputting the input text into the image generation model yields an image that balances differentiated faces and overall image quality, improving the usability of text-based generated images.

[0057] In one embodiment, the loss function for reinforcement training is obtained as follows: The loss function of the Diffusion-DPO algorithm is used as the global loss of the image and fused with the local perceptual loss based on the face region mask in the image to obtain the loss function for reinforcement training.

[0058] Furthermore, the loss function of the Diffusion-DPO algorithm is used as the global loss of the image and fused with the local perceptual loss based on a face region mask in the image, including: Determine the balance coefficient; Based on the balance coefficient, the global loss and the local perception loss are fused.

[0059] Specifically, the reinforcement learning in this application employs the DPO algorithm as its core optimization architecture. The DPO algorithm is a reinforcement learning method directly based on sample preferences. Instead of explicitly building a reward model, it achieves preference-oriented optimization by comparing the generation probabilities of preferred and unpreferred samples. Its basic idea is to maximize the model's log-likelihood on preferred samples while minimizing its log-likelihood on unpreferred samples, ensuring that the model is more inclined to generate results that align with human preferences in the probability space.

[0060] The loss function corresponding to the optimization objective of the DPO algorithm can be expressed as: ; in, This is the loss value. For preference sample pairs (where c represents the input text i.e., prompt, x...), ... + x represents a positive sample image. - (representing negative sample images), D is a dataset consisting of preferred sample pairs. As a reference model, For the sigmoid function, To control the intensity of preference, π θ The model to be optimized (i.e., the pre-trained diffusion model in this application, with parameter θ). This is the expected value.

[0061] The diffusion model generates images through a stepwise denoising process, making it impossible to directly obtain conditional probabilities. , where x represents a sample image pair.

[0062] The Diffusion-DPO algorithm calculates the model's optimal sample at each diffusion time step t. Compared with inferior samples Noise prediction error Furthermore, by applying a preference comparison function in logistic form as a constraint, the following loss function corresponding to the optimization objective is obtained: ; ; in, For the overall loss, The noise predicted by the model at time step t. This is real noise. The weights are for each time step.

[0063] Reference It can be obtained , , as well as The following And so on.

[0064] This loss function realizes the idea of ​​"direct preference optimization" by explicitly comparing the reconstruction error difference between good and bad samples, enabling the model to learn the relative ranking of quality directly from the preference data without having to explicitly construct a reward model.

[0065] However, the Diffusion-DPO algorithm mentioned above usually treats the entire image space equally when calculating the preference loss, without taking into account the differences in importance of different regions to the generated result.

[0066] In face differentiation generation tasks, it is often desirable to optimize the model only in key regions (such as key facial regions). To address this, this application proposes a region-aware diffusion preference optimization algorithm, which introduces a region mask on top of the original Diffusion-DPO algorithm. (H and W represent the height and width of the region mask, respectively), noise prediction error is calculated only in the specified region, thus achieving localized optimization. Therefore, the loss of the optimization objective of region-aware DPO (i.e., local perception loss) can be expressed as: ; ; in, For local perceptual loss, M represents the sum of pixels in any region of the mask. ij This represents the pixel value in the i-th row and j-th column of any region mask.

[0067] By employing a region-aware reinforcement learning strategy, the model optimizes the noise prediction error difference between good and bad samples only within the masked area, thereby achieving differentiated guidance for key regions and promoting the model to learn the feature differences of face regions more accurately.

[0068] Furthermore, traditional reinforcement learning optimization often takes global loss as the sole objective, which can lead to insufficient optimization of the model in key local regions; while simple local optimization may destroy the overall style and structure of the image.

[0069] To enhance facial differentiation while maintaining basic image generation capabilities, this application proposes a weighted balancing mechanism for global and local losses. The loss function of the Diffusion-DPO algorithm is used as the global loss of the image and fused with the region-aware loss. A balancing coefficient is introduced to achieve coordinated optimization of global and local losses.

[0070] Specifically, a balance coefficient with a value range of [0,1] can be determined to adjust the model's attention to the global and local regional structures. In practical applications, this coefficient can be a fixed value or dynamically adjusted using an adaptive strategy.

[0071] This application proposes a balance coefficient based on the proportion of the face region. An adaptive adjustment strategy is defined as follows: ; in, The area of ​​the face mask. This represents the total area of ​​the image.

[0072] This adaptive strategy can dynamically adjust the weights of the local loss based on the proportion of the face region. When the face region accounts for a large proportion of the image or there are many people, the strategy will be more effective. A larger value enhances optimization of local face regions; when the image background is complex or the number of people is small. The value is relatively small to maintain the stability and consistency of the overall generation.

[0073] Furthermore, based on the balance coefficient, the global loss and the local perceptual loss can be fused in the following way to obtain the loss function for reinforcement training: ; in, The loss function is used to enhance training.

[0074] Therefore, this application can achieve a dynamic balance between global and local aspects in multi-person image generation tasks, enabling the model to achieve the optimal trade-off between differentiated face optimization and overall image quality maintenance, and realize a dynamic balance between global structure and local differences, significantly improving the realism and diversity of the generated image.

[0075] Unlike previous methods that only made global adjustments and could not optimize local areas, this application introduces a region-aware reinforcement learning mechanism, incorporates a face region mask into the loss function, and achieves local weighted optimization, effectively improving facial differentiation in multi-person images. At the same time, it proposes a weighted balance mechanism between global and local losses, and by adaptively adjusting the balance coefficient, the model takes into account both local differences and overall consistency, fundamentally solving the problem of facial feature convergence in existing text-based image models.

[0076] Figure 3 This is a schematic diagram of the overall technical solution of the image generation method provided in the embodiments of this application, such as... Figure 3 As shown, this application first constructs pairs of superior and inferior samples (including preferred and inferior samples) of multi-subject images at the data level, using the feature differences of the face region as the reinforcement learning preference signal; then, during the training phase, based on the DPO architecture, it introduces mask information of the face region, enabling the model to focus on key facial regions when learning preferences, thereby achieving region-level reinforcement adjustment; finally, it introduces a balance coefficient-based... The adaptive weighted balancing mechanism of global perception loss (i.e., global loss) and local perception loss achieves a balance between global image consistency and local face differentiation, thereby achieving high-fidelity and significantly different multi-person image generation effects, while effectively avoiding model collapse during continuous optimization.

[0077] The image generation apparatus provided in this application is described below. The image generation apparatus described below can be referred to in correspondence with the image generation method described above.

[0078] Furthermore, this application also provides an image generation apparatus.

[0079] The image generation device includes: The acquisition module is used to acquire the input text; The generation module is used to input the input text into the image generation model and obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0080] The image generation apparatus of this application constructs sample data including sample text, sample image pairs generated from the sample text, and masks of facial regions within the sample image pairs. Since the sample image pairs contain negative sample images generated from the sample text and positive sample images generated from the negative sample images, and provide masks of facial regions within the sample image pairs, it can simultaneously provide differential preference signals and region localization information at the sample level. This allows the model to focus on differential learning of facial features during reinforcement training. Furthermore, when reinforcing the pre-trained diffusion model based on the sample data, the facial region mask is introduced to calculate local perceptual loss, thereby improving facial differentiation in multi-person images. The local perceptual loss and global loss are fused as the loss function for reinforcement training, enabling the model to balance local differences and overall consistency. Consequently, the trained image generation model can balance differentiated faces and overall image quality during image generation. Therefore, after obtaining the input text, inputting the input text into the image generation model yields an image that balances differentiated faces and overall image quality, improving the usability of text-based generated images.

[0081] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute the following method: acquiring input text; The input text is fed into the image generation model to obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0082] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0083] In another aspect, embodiments of this application also provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments, such as: acquiring input text; The input text is fed into the image generation model to obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0084] In another aspect, embodiments of this application also provide a computer program product having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments, such as: acquiring input text; The input text is fed into the image generation model to obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

[0085] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate this application and are not intended to limit this application. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that various combinations, modifications, or equivalent substitutions of the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application.

Claims

1. An image generation method, characterized in that, include: Get the input text; The input text is fed into the image generation model to obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

2. The image generation method according to claim 1, characterized in that, The loss function for the reinforcement training is obtained in the following way: The loss function of the Diffusion-DPO algorithm, which is used as the global loss of the image, is fused with the local perceptual loss based on the face region mask in the image to obtain the loss function for reinforcement training.

3. The image generation method according to claim 2, characterized in that, The local perceptual loss is obtained by introducing a face region mask in the image based on the global loss.

4. The image generation method according to claim 2, characterized in that, The step of fusing the loss function of the Diffusion-DPO algorithm as the global loss of the image with the local perceptual loss based on a face region mask in the image includes: Determine the balance coefficient; Based on the balance coefficient, the global loss and the local perception loss are fused.

5. The image generation method according to claim 4, characterized in that, The balance coefficient is determined based on the proportion of the face region in the image.

6. The image generation method according to any one of claims 1 to 5, characterized in that, The positive sample image is obtained by performing face swapping on the negative sample image.

7. An image generation apparatus, characterized in that, include: The acquisition module is used to acquire the input text; The generation module is used to input the input text into the image generation model and obtain the image output by the image generation model; The image generation model is obtained by reinforcing a pre-trained diffusion model based on sample data; any sample data includes at least one of sample text, a sample image pair generated based on the sample text, and a mask of the face region in the sample image pair; the sample image pair includes a negative sample image generated based on the sample text and a positive sample image generated based on the negative sample image; the loss function of the reinforcement training integrates the global loss of the image and the local perceptual loss based on the mask of the face region in the image.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image generation method as described in any one of claims 1 to 6.

9. A storage medium, said storage medium being a non-transitory computer-readable storage medium, wherein a computer program is stored thereon, characterized in that, When executed by a processor, the computer program implements the image generation method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image generation method according to any one of claims 1 to 6.