Method and apparatus for training image generation model, and device and storage medium

By acquiring the training dataset and semantic prompt text, and combining adversarial loss and semantic control loss to adjust the parameters of the image generation model, the problem of mismatch between generated images and text in the image generation model was solved, and the accuracy of semantic control and training effect were improved.

WO2026152902A1PCT designated stage Publication Date: 2026-07-23PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2025-11-28
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing image generation models are prone to generating images that do not match the objects indicated by the text, resulting in poor semantic control accuracy.

Method used

By acquiring the training dataset, semantic prompt text, and image annotation text, an image generation model is used to perform image semantic control processing. The model parameters are adjusted by combining adversarial loss and semantic control loss to ensure that the generated image matches the object indicated by the semantic prompt text.

Benefits of technology

It improves the semantic control accuracy of the image generation model during the generation process and enhances the training effect of the image generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025138657_23072026_PF_FP_ABST
    Figure CN2025138657_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence model training. Provided are a method and apparatus for training an image generation model, and a device and a storage medium. The method comprises: acquiring a training data set, wherein the training data set comprises a plurality of first images and image annotation text corresponding to each first image; acquiring semantic prompt text; on the basis of an image generation model, the semantic prompt text and the image annotation text, performing image semantic control processing on the first images, so as to obtain second images corresponding to the first images; determining an adversarial loss between the first images and the second images; determining a semantic control loss between the semantic prompt text and the second images; and on the basis of the adversarial loss and the semantic control loss, adjusting a model parameter of the image generation model, so as to improve the accuracy of semantic control of the image generation model during image generation, thereby improving the training effect of the image generation model. The image generation model can be applied to the field of financial technology, the field of healthcare, etc., so as to generate corresponding images.
Need to check novelty before this filing date? Find Prior Art

Description

Training methods, devices, equipment, and storage media for image generation models

[0001] This application claims priority to Chinese Patent Application No. 2025100668707, filed on January 15, 2025, entitled “Training Method, Apparatus, Device and Storage Medium for Image Generation Model”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence model training technology, and in particular to a training method, apparatus, device and storage medium for an image generation model. Background Technology

[0003] Currently, traditional techniques can generate images using image generation models, such as the Stable Diffusion (SD) model, combined with corresponding text. However, the inventors realized that during the image generation process using image generation models, it is easy for the objects contained in the generated image to not match the objects indicated by the text, resulting in poor semantic control accuracy of the image generation model during image generation. Summary of the Invention

[0004] The main objective of this application is to provide a training method, apparatus, device, and storage medium for an image generation model, aiming to improve the semantic control accuracy of the image generation model during the image generation process, thereby enhancing the training effect of the image generation model.

[0005] Firstly, this application provides a method for training an image generation model, comprising:

[0006] Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image;

[0007] Obtain semantic prompt text;

[0008] Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image;

[0009] Based on the first image and the second image, determine the adversarial loss between the first image and the second image;

[0010] Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image;

[0011] The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

[0012] Secondly, this application also provides a training apparatus for an image generation model, the training apparatus comprising:

[0013] The first acquisition module is used to acquire a training dataset, which includes multiple first images and image annotation text corresponding to each first image.

[0014] The second acquisition module is used to acquire semantic prompt text;

[0015] The processing module is used to perform image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image;

[0016] The first loss determination module is used to determine the adversarial loss between the first image and the second image based on the first image and the second image;

[0017] The second loss determination module is used to determine the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image;

[0018] The model training module is used to adjust the model parameters of the image generation model based on the adversarial loss and the semantic control loss.

[0019] Thirdly, this application also provides a computer device, which includes a memory and a processor;

[0020] The memory is used to store computer programs;

[0021] The processor is configured to execute the computer program and, in executing the computer program, perform the following steps:

[0022] Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image;

[0023] Obtain semantic prompt text;

[0024] Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image;

[0025] Based on the first image and the second image, determine the adversarial loss between the first image and the second image;

[0026] Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image;

[0027] The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

[0028] Fourthly, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following steps:

[0029] Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image;

[0030] Obtain semantic prompt text;

[0031] Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image;

[0032] Based on the first image and the second image, determine the adversarial loss between the first image and the second image;

[0033] Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image;

[0034] The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

[0035] This application provides a training method, apparatus, device, and storage medium for an image generation model. The training method includes: acquiring a training dataset, which includes multiple first images and corresponding image annotation text for each first image; acquiring semantic prompt text; performing image semantic control processing on the first images based on the image generation model, the semantic prompt text, and the image annotation text to obtain a second image corresponding to the first image; determining the adversarial loss between the first image and the second image based on the first image and the second image; determining the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image; and adjusting the model parameters of the image generation model based on the adversarial loss and the semantic control loss.

[0036] For example, during the training of an image generation model, the model parameters can be adjusted based on adversarial loss and semantic control loss. Since the adversarial loss is determined based on the first and second images, the image generation model can use this loss to evaluate whether the second image is similar to the first image. Similarly, since the semantic control loss is determined based on the semantic prompt text and the second image, the image generation model can use this loss to evaluate whether the objects contained in the second image match the objects indicated by the semantic prompt text. Therefore, by adjusting the model parameters based on the adversarial and semantic control losses, the image generation model can perform semantic control processing on the first image based on the semantic prompt text to obtain a second image that is similar to the first image and contains objects that match the objects indicated by the semantic prompt text. This improves the semantic control accuracy of the image generation model during image generation, thereby enhancing the training effect of the image generation model.

[0037] The trained image generation model obtained through this training method can be applied to fields such as fintech and healthcare. In some implementations, for auto insurance in the fintech sector, responding to the marketing needs of auto insurance products—for example, the need to generate a large number of car-related marketing images—the trained image generation model can be used to generate a large number of second images similar to the first image, containing objects that match the objects indicated by the semantic prompts, by combining a first image provided by the auto insurance business with semantic prompt text related to auto insurance product marketing. Based on the generation of these second images, which can be used for marketing auto insurance products, the convenience of marketing these products is improved. In other implementations, for medical insurance in the healthcare sector, responding to the marketing needs of medical insurance products—for example, the need to generate a large number of medical insurance product marketing images—the trained image generation model can be used to generate a large number of second images similar to the first image, containing objects that match the objects indicated by the semantic prompts, by combining a first image provided by the medical insurance business with semantic prompt text related to medical insurance product marketing. Based on the generation of these second images, which can be used for marketing medical insurance products, the convenience of marketing these products is improved. Of course, it is not limited to this, and no restrictions are set here. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 is a flowchart illustrating a training method for an image generation model provided in an embodiment of this application;

[0040] Figure 2 is a schematic flowchart of an image generation method according to an embodiment of this application;

[0041] Figure 3 is a schematic block diagram of a training device for an image generation model provided in an embodiment of this application;

[0042] Figure 4 is a schematic block diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0045] This application provides a training method, apparatus, device, and storage medium for an image generation model. The training method for this image generation model can be applied to a computer device, such as a tablet, laptop, or desktop computer. It can also be applied to a server, which can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0046] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0047] Please refer to Figure 1, which is a flowchart illustrating a training method for an image generation model according to an embodiment of this application. It should be noted that the training method for the image generation model provided in this embodiment can be used on a computer device, and of course, on a server. For example, the server can obtain a training dataset from the computer device and process the training dataset according to the training method of the image generation model to obtain a second image corresponding to a first image. Then, it determines the adversarial loss between the first and second images and the semantic control loss between the semantic prompt text and the second image, thereby adjusting the model parameters of the image generation model based on the adversarial loss and the semantic control loss. For example, the server can send the trained image generation model obtained according to the training method of the image generation model to the computer device, so that the computer device can generate the corresponding second image based on the image generation model; however, it is not limited to this, and the trained image generation model can be applied to different fields such as fintech and healthcare to generate corresponding images, without limitation.

[0048] In practice, computer equipment includes, but is not limited to, any of the following: tablet computers, laptop computers, and desktop computers; the server can be a single server, a server cluster, or a cloud server that provides cloud computing services.

[0049] As shown in Figure 1, the training method of this image generation model includes steps S101 to S106.

[0050] Step S101: Obtain the training dataset, which includes multiple first images and the corresponding image annotation text for each first image.

[0051] For example, a computer device can acquire multiple first images to construct a training dataset. This training dataset can be used by the computer device to train an image generation model. If the training dataset includes first images, the computer device can input the first images into the image generation model, which can then use them as the basis for generating subsequent second images corresponding to those first images. The objects included in different first images can be the same or different; this is not limited. Objects in the first images can include text, objects, etc., without limitation. Text includes, for example, artistic fonts, text in preset fonts, letters, numbers, etc., without limitation. Objects include, for example, cars, appliances, household items, etc., without limitation.

[0052] In one exemplary embodiment, when training the image generation model, the computer device can acquire images related to different fields such as fintech and healthcare, based on the consideration that the trained image generation model can be applied to various fields such as fintech and healthcare, and use these images as multiple first images included in the training dataset. Images related to the fintech field include, for example, vehicle images involved in auto insurance business, auto insurance scene images, and claims images related to claims information involved in insurance claims business, etc., without limitation. Images related to the healthcare field include, for example, images related to different types of diseases, images related to different types of disease symptoms, etc., without limitation. Of course, during subsequent training, the image generation model can also be trained based on images unrelated to different fields such as fintech and healthcare to train how to generate images related to different fields such as fintech and healthcare according to the instructions of semantic prompt text, without limitation.

[0053] Given multiple first images, a computer device can determine the image annotation text corresponding to each first image, and combine the first images and their corresponding image annotation texts to construct a training dataset. The image annotation text can be used to indicate the objects included in the first images. For example, the computer device can use a pre-defined image annotation tool to perform image annotation processing on the first images, obtaining the image annotation text corresponding to each first image. The image annotation tool can include multimodal visual-text large language models, such as the BLIP2 model, etc., and is not limited here. For example, the image annotation tool can recognize text, objects, etc., in the first images, obtaining the corresponding recognition results. Accordingly, the image annotation tool can set the recognition result as the image annotation text corresponding to the first image in text form.

[0054] Once the training dataset is obtained, the computer device can then train the image generation model on how to improve the semantic control accuracy during the image generation process, based on the multiple first images included in the training dataset and the corresponding image annotation text for each first image.

[0055] In some implementations, multiple preset images are acquired; from the multiple preset images, preset images that do not meet the preset image filtering conditions are removed to obtain a first image; image annotation processing is performed on the first image to obtain the image annotation text corresponding to the first image.

[0056] In training an image generation model to improve its semantic control accuracy during image generation, a computer device can acquire as many first images as possible for the model to use. Based on this, the computer device can acquire multiple preset images to filter and / or process to obtain the first image. For example, the computer device can acquire multiple preset images based on publicly available data, such as open-source datasets, websites, etc. Accordingly, the computer device can determine whether any preset images do not meet the preset image filtering criteria, and then remove these to obtain the first image. Preset image filtering criteria may include at least one of the following: the presence of a watermark or copyright ownership. However, preset image filtering criteria are not limited to these; they can be pre-set or user-defined, and are not restricted here.

[0057] Once the first image is acquired, the computer device can determine the image annotation text corresponding to the first image. For example, by using a preset image annotation tool, the first image can be processed to obtain the image annotation text corresponding to the first image.

[0058] In this way, the computer device can select the first image from multiple preset images based on preset image selection criteria, which improves the ease of determining the first image. Correspondingly, the first image can be used by the computer device for image annotation processing to determine the corresponding image annotation text. The image annotation text can be used to indicate the objects included in the first image, which can then be used to train the image generation model.

[0059] For example, after acquiring the first image, the computer device can process the first image accordingly to acquire a larger number of first images using a limited number of first images, for subsequent training of the image generation model.

[0060] For example, the first image can be resized to obtain the first image at different sizes.

[0061] The computer device can perform image resizing processing on the first image. For example, the computer device can transform the first image into a first image of different sizes. Different image sizes include, for example, 1024 pixels × 1024 pixels, 768 pixels × 768 pixels, 1024 pixels × 768 pixels, etc., and are not limited here. Different image sizes can be preset or set by the user, and are not limited here. When the first image can be transformed to obtain first images of different sizes, the number of first images of different sizes can be greater than the number of original first images, and all first images of different sizes can be used to train the image generation model. This is beneficial to increasing the amount of training data for the image generation model, thereby improving the training effect of the image generation model.

[0062] For example, image annotation generation is performed on the first image at different image sizes to obtain the corresponding image annotation text for the first image at different image sizes.

[0063] Given first images at different image sizes, the computer device can determine the corresponding image annotation text for each image at different sizes. For example, using a pre-defined image annotation tool, image annotation processing can be performed on the first images at different sizes to obtain the corresponding image annotation text for each image at different sizes. The image annotation text can be used to indicate the objects included in the first images at different sizes, for subsequent training of the image generation model.

[0064] Step S102: Obtain semantic prompt text.

[0065] For example, during the training of an image generation model, a computer device can acquire semantic prompt text. For instance, the semantic prompt text could be input by a relevant person. When the computer device detects text input by this person, it can acquire the corresponding input text. If the input text does not meet preset text filtering criteria, it can be discarded. If the input text meets the preset text filtering criteria, it can be identified as semantic prompt text. Preset text filtering criteria may include that the input text is not garbled text, or that the input text is not purely punctuation, etc., and are not limited here. Preset text filtering criteria can be pre-set or user-defined, and are not limited here.

[0066] Semantic cue text can be used by the image generation model to determine the semantic control requirements when performing semantic control processing on the first image. The semantic cue text can be used to indicate changes to objects included in the first image. Changes to objects included in the first image may include, for example, adding new objects to the first image and / or deleting old objects from the first image.

[0067] Take auto insurance in the fintech sector as an example. Responding to the marketing needs of auto insurance products, the business requires image generation models to provide a large number of marketing images. Computer equipment can then determine semantic prompt text based on objects related to auto insurance products. These objects include, for example, customer categories, types of vehicles covered by the insurance, insurance types, and insurance scenarios. The computer equipment can then use these customer categories, vehicle types covered by the insurance, insurance types, and insurance scenarios as semantic prompt text to input into the image generation model.

[0068] Take the health insurance business in the healthcare field as an example. Responding to the marketing needs of health insurance products, the business requires an image generation model that can provide a large number of marketing images for these products. The computer device can then determine semantic prompt text based on objects related to the health insurance products. These related objects include, for example, the customer category, the types of diseases covered, and the type of health insurance. The computer device can then use these factors as semantic prompt text to input into the image generation model.

[0069] The semantic prompt text can be combined with the image generation model to perform image semantic control processing on the first image to obtain the second image corresponding to the first image.

[0070] Thus, with semantic prompt text obtained, the image generation model can use the semantic prompt text as a basis to perform image semantic control processing on the first image, so as to determine the second image corresponding to the first image in the future, which is conducive to improving the convenience of subsequent image semantic control processing of the first image.

[0071] Step S103: Based on the image generation model, semantic prompt text, and image annotation text, perform image semantic control processing on the first image to obtain the second image corresponding to the first image.

[0072] When a computer device acquires a first image and semantic prompt text, it can input the first image and semantic prompt text into an image generation model. Since the first image can carry corresponding image annotation text, the corresponding image annotation text can also be input into the image generation model. Upon receiving the first image, semantic prompt text, and corresponding image annotation text, the image generation model can combine the semantic prompt text and the image annotation text to perform image semantic control processing on the first image, thereby obtaining a second image corresponding to the first image.

[0073] In some implementations, the image generation model can perform image semantic control processing on the first image based on the semantic information of the semantic prompt text. The image annotation text can provide additional semantic information for the first image. The image generation model can determine the semantic information of both the semantic prompt text and the image annotation text. For example, the image generation model can use natural language processing techniques, such as word embedding and semantic analysis, to convert the semantic prompt text and the image annotation text into representations that the image generation model can understand, thereby obtaining the semantic information of each text.

[0074] Accordingly, during the semantic control processing of the first image, the image generation model can determine whether the first image covers the object indicated by the semantic prompt text based on the semantic information of the first image and its corresponding image annotation text. It can then further modify the object in the first image by combining the semantic information of the semantic prompt text. For example, if, based on the semantic information of the first image and its image annotation text, it is determined that the first image does not cover object 1 indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that object 1 needs to be added to the first image, the image generation model can add object 1 to the first image to determine the second image corresponding to the first image. As another example, if, based on the semantic information of the first image and its image annotation text, it is determined that the first image covers object 2 indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that object 2 needs to be deleted from the first image, the image generation model can delete object 2 from the first image to determine the second image corresponding to the first image. Of course, this is not a limitation and is not set forth here.

[0075] Thus, based on the image generation model, semantic prompt text, and image annotation text, the image semantic control processing is performed on the first image to obtain the second image corresponding to the first image. The image generation model can be trained to improve the accuracy of semantic control in the image generation process so that the second image is similar to the first image and the objects contained in the second image match the objects indicated by the semantic prompt text, thereby improving the training effect of the image generation model.

[0076] For example, the image generation model may include a text / image encoder and a stable diffusion model with a semantic fine-tuning subnetwork. In one exemplary implementation, the text / image encoder may use OpenClip-G. The stable diffusion model may use SDXL1.0. The semantic fine-tuning subnetwork may include a LoRA subnetwork. The LoRA subnetwork may be set in the attention structure of the stable diffusion network. The attention structure of the stable diffusion subnetwork may be set in a U-Net architecture. The U-Net architecture includes, for example, convolutional layers, attention layers, upsampling layers, and downsampling layers, and the LoRA subnetwork can be set in one of the attention layers of the U-Net architecture. However, this is not a limitation and is not intended to restrict the implementation.

[0077] In some implementations, a text / image encoder based on an image generation model performs embedding feature extraction processing on semantic prompt text, image annotation text, and a first image, respectively, to obtain a first text embedding feature corresponding to the semantic prompt text, a second text embedding feature corresponding to the image annotation text, and a first image embedding feature corresponding to the first image; a stable diffusion network based on an image generation model and a semantic fine-tuning sub-network set in the stable diffusion network performs image semantic control processing on the first image embedding feature according to the first text embedding feature and the second text embedding feature, to obtain a second image corresponding to the first image.

[0078] For example, an image generation model can input the acquired semantic prompt text, image annotation text, and a first image into a text / image encoder. The text / image encoder can perform embedding feature extraction on the text and image, allowing it to extract embedding features from the semantic prompt text, image annotation text, and the first image respectively, resulting in a first text embedding feature corresponding to the semantic prompt text, a second text embedding feature corresponding to the image annotation text, and a first image embedding feature corresponding to the first image. The first and second text embedding features allow the image generation model to determine whether the first image covers the object indicated by the semantic prompt text, providing subsequent instructions for semantic control processing of the first image. The first image embedding feature allows the image generation model to perform corresponding semantic control processing, thereby determining the second image corresponding to the first image.

[0079] Given a first text embedding feature, a second text embedding feature, and a first image embedding feature, the image generation model can input these features into a stable diffusion network. The stable diffusion network can generate images based on text, and a semantic fine-tuning subnetwork within it can fine-tune the network. Therefore, the stable diffusion network with the semantic fine-tuning subnetwork allows the image generation model to perform image semantic control processing on the first image embedding feature based on the first and second text embedding features, thereby obtaining a second image corresponding to the first image. For example, the stable diffusion network with the semantic fine-tuning subnetwork can perform image semantic control processing on the first image embedding feature based on the semantic information of the first and second text embedding features, changing the objects included in the first image in response to the instruction of the semantic prompt text, so that the second image matches the semantic information of the semantic prompt text.

[0080] Taking auto insurance in the fintech field as an example. The first image may include at least one of vehicle images and non-vehicle images. The image annotation text corresponding to the first image can be used to indicate that the objects in the vehicle image include vehicles, while the objects in the non-vehicle image do not include vehicles. The objects indicated by the semantic prompt text may include, for example, auto insurance customer categories, types of vehicles covered by auto insurance, auto insurance types, and auto insurance scenarios. The semantic information of the semantic prompt text is used to indicate the objects indicated by the semantic prompt text in the first image. The image generation model can use a text / image editor to perform embedding feature extraction processing on the image annotation text and semantic prompt text corresponding to each of the vehicle image, non-vehicle image, vehicle image, and non-vehicle image to determine the first image embedding feature corresponding to each of the vehicle image and non-vehicle image, the first text embedding feature corresponding to the semantic prompt text, and the second text embedding feature corresponding to the image annotation text. The image generation model can determine whether objects in vehicle images and non-vehicle images include car insurance customer categories, car insurance covered vehicle types, car insurance types, and car insurance scenarios based on first and second text embedding features, respectively. When the corresponding objects are not covered in the vehicle images and / or non-vehicle images, a stable diffusion network with a semantic fine-tuning sub-network is used, combined with the first and second text embedding features, to perform image semantic control processing on the vehicle images and non-vehicle images respectively, thereby obtaining second images corresponding to each vehicle image and non-vehicle image. The objects in the second images corresponding to each vehicle image and non-vehicle image can be supplemented with at least one of the following: car insurance customer category, car insurance covered vehicle type, car insurance type, and car insurance scenario, according to the semantic information indicated by the semantic prompt text. However, this is not a limitation and is not set forth herein.

[0081] Taking medical insurance in the healthcare field as an example. The first image may include at least one of disease images and non-disease images. The image annotation text corresponding to the first image can be used to indicate that the objects in the disease image include the corresponding disease type, while the objects in the non-disease image do not include any disease type. The objects indicated by the semantic prompt text may include, for example, the disease type covered by medical insurance, the medical insurance customer category, and the type of medical insurance. The semantic information of the semantic prompt text is used to indicate the objects indicated by the semantic prompt text in the first image. The image generation model can use a text / image editor to perform embedding feature extraction processing on the image annotation text and semantic prompt text corresponding to the disease image, non-disease image, and disease image and non-disease image respectively, to determine the first image embedding feature corresponding to the disease image and non-disease image, the first text embedding feature corresponding to the semantic prompt text, and the second text embedding feature corresponding to the image annotation text. The image generation model can determine whether objects in disease images and non-disease images include disease types covered by medical insurance, medical insurance customer categories, and medical insurance types, based on first and second text embedding features, respectively. When the disease image and / or non-disease image do not cover the corresponding objects, a stable diffusion network with a semantic fine-tuning sub-network, combined with the first and second text embedding features, performs semantic control processing on the disease image and non-disease image respectively, thereby obtaining the corresponding second images for each. The objects in the corresponding second images for both disease and non-disease images can be supplemented with at least one of the following: disease type covered by medical insurance, medical insurance customer category, and medical insurance type, according to the semantic information indicated by the semantic prompt text. However, this is not a limitation and is not set forth herein.

[0082] In this way, the image generation model can determine the first text embedding feature corresponding to the semantic prompt text, the second text embedding feature corresponding to the image annotation text, and the first image embedding feature corresponding to the first image through a text / image encoder. Then, through a stable diffusion network with a semantic fine-tuning subnetwork, the first image embedding feature is processed for image semantic control based on the first and second text embedding features to obtain the second image corresponding to the first image. This allows the image generation model to determine whether the object indicated by the semantic prompt text matches the object indicated by the image annotation text based on the first and second text embedding features. If there is a discrepancy, the image generation model can perform image semantic control processing on the first image based on the object indicated by the semantic prompt text, thereby matching the semantic information of the second image corresponding to the first image with that of the semantic prompt text. This is beneficial for training the image generation model to improve the accuracy of semantic control of images during image generation, thus improving the training effect of the image generation model.

[0083] Step S104: Determine the adversarial loss between the first image and the second image based on the first image and the second image.

[0084] For example, when training an image generation model to improve the accuracy of semantic control over images during image generation, a computer device can introduce the adversarial mechanism of a Generative Adversarial Network (GAN) into the image generation model. This GAN mechanism is used to determine the adversarial loss between the first and second images. For instance, when determining the adversarial loss, the GAN needs to determine the loss of the generator and the discriminator. The generator's goal is to generate realistic fake images to deceive the discriminator; therefore, the generator's loss can be associated with the probability that the discriminator classifies a fake image as a real image. The discriminator's goal is to accurately distinguish between real and fake images; therefore, the discriminator's loss can be associated with the probability that the discriminator classifies a real image as real and the probability that it classifies a fake image as fake. Based on this, the image generation model can set the first image as a real image involved in the GAN network and the second image as a fake image involved in the GAN network, allowing the image generation model to determine the adversarial loss between the first and second images. For example, an image generation model can assess the probability of identifying the first image as the second image by determining whether the first image and the second image are similar, thereby determining the adversarial loss.

[0085] In the process of determining whether a first image and a second image are similar, the image generation model can assess the probability of identifying the first image as the second image by judging the image similarity between the first image and the second image, as well as the semantic similarity between the first image and the second image, thereby determining the adversarial loss between the first image and the second image.

[0086] In some implementations, a perceptual feature extraction network based on an image generation model performs perceptual feature extraction processing on a first image and a second image respectively, to obtain a first perceptual feature corresponding to the first image and a second perceptual feature corresponding to the second image; determines the perceptual loss between the first and second perceptual features; a residual network based on an image generation model performs semantic classification processing on the first image and the second image respectively, to obtain a first classification result corresponding to the first image and a second classification result corresponding to the second image; determines the classification loss between the first and second classification results; and determines the adversarial loss based on the sum of the perceptual loss and the classification loss.

[0087] Image generation models can incorporate perceptual feature extraction networks. These networks extract perceptual features from images to obtain their corresponding perceptual characteristics. Perceptual features indicate aspects of an image that can be perceived and recognized by the human visual system. These features can be related to the visual content of the image and may include, for example, color, texture, shape, edges, and higher-level semantic information such as objects and scenes; this is not a limitation. Perceptual feature extraction networks may include, for example, deep learning networks such as the VGG (Visual Geometry Group) network; this is also not a limitation.

[0088] For example, an image generation model can input a first image and a second image into a perceptual feature extraction network, so that the perceptual feature extraction network can perform perceptual feature extraction processing on the first image and the second image respectively, to obtain the first perceptual feature corresponding to the first image and the second perceptual feature corresponding to the second image.

[0089] Accordingly, the first and second perceptual features can be used by the image generation model to determine the image similarity between the first and second images. For example, the image generation model can determine the difference between the first and second perceptual features to determine the perceptual loss between them. The perceptual loss can be positively correlated with the image similarity between the first and second images. The smaller the perceptual loss, the higher the image similarity between the first and second images; conversely, the larger the perceptual loss, the lower the image similarity between the first and second images.

[0090] Image generation models can incorporate residual networks (ResNet). ResNet can be used for semantic classification of images, yielding classification results. These results can indicate the semantic information of the image. The classification results can be related to the semantic information of the image, and may include, for example, object classification results, scene classification results, etc., without limitation.

[0091] For example, an image generation model can input a first image and a second image into a residual network, so that the residual network can perform semantic classification processing on the first image and the second image respectively, to obtain a first classification result corresponding to the first image and a second classification result corresponding to the second image.

[0092] Accordingly, the first and second classification results can be used by the image generation model to determine the semantic similarity between the first and second images. For example, the image generation model can determine the differences between the first and second classification results to determine whether they belong to the same category, and thus determine the classification loss between the first and second classification results. The classification loss can be minimized if the first and second classification results belong to the same category; the smaller the classification loss, the higher the semantic similarity between the first and second images; conversely, the larger the classification loss, the lower the semantic similarity between the first and second images.

[0093] Given the perceptual loss between the first and second perceptual features, and the classification loss between the first and second classification results, the image generation model can determine the adversarial loss based on the sum of the perceptual and classification losses. For example, during training, the image generation model can adjust its parameters based on a strategy of minimizing the adversarial loss, i.e., minimizing the sum of the perceptual and classification losses, so that the second image generated by the image generation model is similar to the first image generated from the input image.

[0094] In this way, the image generation model can determine its adversarial loss based on the first and second images. This adversarial loss can be used to train the image generation model further, increasing the similarity between the second image generated by the model and the first image from the input image generation model, thus improving the training performance of the image generation model.

[0095] Step S105: Determine the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image.

[0096] For example, when training an image generation model to improve the accuracy of semantic control over an image during the image generation process, a computer device can determine whether the semantic information of the second image matches the semantic information of the semantic prompt text based on the semantic prompt text and the second image, so as to determine the accuracy of image semantic control when the image generation model performs image semantic control processing on the first image in response to the instruction of the semantic prompt text.

[0097] In some implementations, a text / image encoder based on an image generation model performs embedding feature extraction processing on the semantic prompt text and the second image, respectively, to obtain the first text embedding feature corresponding to the semantic prompt text and the second image embedding feature corresponding to the second image; and determines the semantic control loss based on the distance between the first text embedding feature and the second image embedding feature.

[0098] The image generation model can be equipped with a text / image encoder. The text / image encoder can be used to extract embedding features from at least one of the text and the image, obtaining the corresponding embedding features for each of the text and the image. These embedding features can be used to indicate the semantic information of at least one of the text and the image.

[0099] For example, an image generation model can input semantic prompt text and a second image into a text / image encoder, so that the text / image encoder can perform embedding feature extraction processing on the semantic prompt text and the second image respectively, to obtain the first text embedding feature corresponding to the semantic prompt text and the second image embedding feature corresponding to the second image.

[0100] Accordingly, the first text embedding feature and the second image embedding feature can be used by the image generation model to determine whether the semantic information of the second image matches the semantic information of the semantic prompt text. For example, the higher the matching degree between the object in the second image and the object indicated by the semantic prompt text, the higher the probability that the semantic information of the second image matches the semantic information of the semantic prompt text. Based on this, the image generation model can determine the distance between the first text embedding feature and the second image embedding feature to evaluate the matching degree between the object in the second image and the object indicated by the semantic prompt text, thereby obtaining the semantic control loss. For example, the image generation model can determine the semantic control loss by calculating measurable distances such as cosine similarity between the first text embedding feature and the second image embedding feature, without limitation.

[0101] When the semantic control loss is determined based on the distance between the first text embedding features and the second image embedding features, the semantic control accuracy of the image generation model during image generation can be evaluated based on this loss. A larger semantic control loss indicates poorer semantic control accuracy, while a smaller loss indicates better accuracy. For example, during training, the image generation model's parameters can be adjusted based on a strategy that minimizes the semantic control loss—that is, minimizes the distance between the first text embedding features and the second image embedding features—so that the object indicated by the second image generated by the model matches the object indicated by the semantic prompt text.

[0102] In this way, the image generation model can determine its semantic control loss based on the semantic prompt text and the second image. This semantic control loss can be used to train the image generation model further, improving the matching degree between the object indicated by the second image generated by the model and the object indicated by the semantic prompt text. This is beneficial for improving the semantic control accuracy of the image generation model during image generation, thereby enhancing the training effect of the model.

[0103] Step S106: Adjust the model parameters of the image generation model based on the adversarial loss and semantic control loss.

[0104] Given the adversarial loss and semantic control loss of the image generation model, the model parameters can be adjusted based on these losses. For example, the image generation model can be trained by minimizing both the adversarial and semantic control losses. Consequently, by minimizing both losses, the trained image generation model can improve the semantic control accuracy during the generation of the second image corresponding to the first image, thus enhancing the overall accuracy and quality of the second image generation.

[0105] In some implementations, the preset network parameters of the semantic fine-tuning subnetwork are adjusted based on the adversarial loss and the semantic control loss, while the preset network parameters of the subnetworks in the stable diffusion network other than the semantic fine-tuning subnetwork remain unchanged.

[0106] For example, when training an image generation model to improve its semantic control accuracy, the preset network parameters of the sub-networks in the stable diffusion network, except for the semantic fine-tuning sub-network, can be frozen. Only the preset network parameters of the semantic fine-tuning sub-network can be adjusted and updated. During training, the image generation model can minimize adversarial loss and semantic control by fine-tuning the preset network parameters of the semantic fine-tuning sub-network. This improves the semantic control accuracy of the image generation model in the process of generating the second image, as well as the image generation accuracy and effect of the second image.

[0107] In one exemplary embodiment, the trained image generation model can be used to generate images. As shown in FIG2, the process of the image generation method using the trained image generation model may include steps S201 to S203.

[0108] Step S201: Obtain the first image and the image annotation text corresponding to the first image.

[0109] For example, the first image can be preset or set by the user. For instance, a computer device may have multiple preset images preset. The user can select one or more preset images as the first image, which will then be used to generate the corresponding second image. Alternatively, the user can input one or more preset images they have acquired into the computer device, allowing the computer device to identify the received preset image as the first image. Of course, this is not a limitation and is not set here.

[0110] Once the first image is acquired, the computer device can perform image annotation processing on the first image to obtain the image annotation text corresponding to the first image.

[0111] Taking auto insurance in the fintech field as an example, the first image can include either a vehicle image or a non-vehicle image; there is no limitation on this. If the first image is a vehicle image, the corresponding image annotation text might include "vehicle." If the first image is a non-vehicle image, the corresponding image annotation text might include "no vehicle." However, this is not limited to these limitations. For example, computer devices can also identify whether vehicle and non-vehicle images contain pedestrians, and add text such as "pedestrian" or "no pedestrian" to the corresponding image annotation text. There are no further limitations on this.

[0112] Taking medical insurance in the healthcare field as an example, the first image can include either a disease image or a non-disease image; there is no limitation on this. If the first image is a disease image, the image annotation text corresponding to the disease image may include, for example, the disease type corresponding to the disease image. If the first image is a non-disease image, the image annotation text corresponding to the non-disease image may include, for example, "no corresponding disease type." However, this is not limited to these limitations. For example, computer devices can also identify whether disease images and non-disease images contain environmental regions, and add text such as "environment type" or "no specific environment type" to the corresponding image annotation text. There are no further limitations on this.

[0113] Step S202: Obtain semantic prompt text.

[0114] For example, semantic prompt text can be pre-set or user-defined. For instance, a computer device may have multiple pre-set prompt texts. The user can select one or more pre-set prompt texts as semantic prompt texts for subsequent generation of a corresponding second image. Alternatively, the user can manually input one or more pre-set prompt texts into the computer device, allowing the device to determine that the received prompt text is a semantic prompt text. Of course, this is not a limitation and is not specified here.

[0115] Taking auto insurance in the fintech sector as an example, semantic prompts might include phrases like "Add information about the types of vehicles covered by the auto insurance product, the applicable auto insurance scenarios, and relevant marketing statements for the auto insurance product to the image, and delete the image of people in the image." Of course, semantic prompts are not limited to these examples, and no restrictions are imposed here.

[0116] Taking medical insurance in the healthcare field as an example, semantic prompt text might include phrases like "Add the types of diseases covered by medical insurance to the image, the judgment result on whether medical insurance covers the types of diseases indicated in the image, the applicable customer categories for medical insurance, and add corresponding marketing statements for medical products to the image." Of course, semantic prompt text is not limited to these examples, and no restrictions are imposed here.

[0117] Step S203: Based on the image generation model, semantic prompt text, and image annotation text, perform image semantic control processing on the first image to obtain the second image corresponding to the first image.

[0118] For example, a computer device can input the acquired first image, the corresponding image annotation text, and the semantic prompt text into an image generation model, so that the image generation model responds to the semantic prompt text and the image annotation text to perform image semantic control processing on the first image to obtain the corresponding second image.

[0119] Taking auto insurance in the fintech field as an example, the semantic prompt text includes instructions such as "Add the types of vehicles covered by the auto insurance product, the applicable auto insurance scenarios, and the corresponding marketing statements for the auto insurance product to the image, and delete the image area containing people from the image." If the first image is a vehicle image, the corresponding image annotation text might include, for example, "vehicle" and "pedestrian." The image generation model can then determine, based on the semantic prompt text and the image annotation text, that it needs to add the types of vehicles covered by the auto insurance product, the applicable auto insurance scenarios, and the marketing statements for the auto insurance product to the vehicle image, and delete the image area containing pedestrians from the vehicle image. Similarly, if the first image is not a vehicle image, the corresponding image annotation text might include, for example, "non-vehicle" and "pedestrian." The image generation model can then determine, based on the semantic prompt text and the image annotation text, that it needs to add the vehicle, the types of vehicles covered by the auto insurance product, the applicable auto insurance scenarios, and the marketing statements for the auto insurance product to the non-vehicle image, and delete the image area containing pedestrians from the non-vehicle image. Of course, this is not a limitation and is not set forth here.

[0120] Taking medical insurance in the healthcare field as an example, the semantic prompt text might include instructions such as "Add the types of diseases covered by medical insurance to the image, the judgment result on whether medical insurance covers the types of diseases indicated in the image, the applicable customer categories for medical insurance, and add corresponding marketing statements for medical products to the image." If the first image is a disease image, the image annotation text corresponding to the disease image might include, for example, the disease type corresponding to the disease image. The image generation model can then determine, based on the semantic prompt text and the image annotation text, that it needs to add the types of diseases covered by medical insurance, the judgment result on whether medical insurance covers the types of diseases indicated in the image, the applicable customer categories for medical insurance, and marketing statements for medical products to the disease image. If the first image is a non-disease image, the image annotation text corresponding to the non-disease image might include, for example, no corresponding disease type. The image generation model can then determine, based on the semantic prompt text and the image annotation text, that it needs to add the types of diseases covered by medical insurance, the applicable customer categories for medical insurance, and marketing statements for medical products to the non-disease image. Of course, this is not a limitation and is not set forth here.

[0121] Thus, a well-trained image generation model can be applied to one or more semantic prompt texts. By performing image semantic control processing on the first image, the well-trained image generation model can achieve good semantic control accuracy on the second image corresponding to the first image during the generation process.

[0122] The image generation model training method provided in the above embodiments involves: acquiring a training dataset, which includes multiple first images and corresponding image annotation text for each first image; acquiring semantic prompt text; performing image semantic control processing on the first images based on the image generation model, the semantic prompt text, and the image annotation text to obtain a second image corresponding to the first image; determining the adversarial loss between the first and second images based on the first and second images; determining the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image; and adjusting the model parameters of the image generation model based on the adversarial loss and the semantic control loss.

[0123] For example, during the training of an image generation model, the model parameters can be adjusted based on adversarial loss and semantic control loss. Since the adversarial loss is determined based on the first and second images, the image generation model can use this loss to evaluate whether the second image is similar to the first image. Similarly, since the semantic control loss is determined based on the semantic prompt text and the second image, the image generation model can use this loss to evaluate whether the objects contained in the second image match the objects indicated by the semantic prompt text. Therefore, by adjusting the model parameters based on the adversarial and semantic control losses, the image generation model can perform semantic control processing on the first image based on the semantic prompt text to obtain a second image that is similar to the first image and contains objects that match the objects indicated by the semantic prompt text. This improves the semantic control accuracy of the image generation model during image generation, thereby enhancing the training effect of the image generation model.

[0124] The trained image generation model obtained through this training method can be applied to fields such as fintech and healthcare. In some implementations, for auto insurance in the fintech sector, responding to the marketing needs of auto insurance products—for example, the need to generate a large number of car-related marketing images—the trained image generation model can be used to generate a large number of second images similar to the first image, containing objects that match the objects indicated by the semantic prompts, by combining a first image provided by the auto insurance business with semantic prompt text related to auto insurance product marketing. Based on the generation of these second images, which can be used for marketing auto insurance products, the convenience of marketing these products is improved. In other implementations, for medical insurance in the healthcare sector, responding to the marketing needs of medical insurance products—for example, the need to generate a large number of medical insurance product marketing images—the trained image generation model can be used to generate a large number of second images similar to the first image, containing objects that match the objects indicated by the semantic prompts, by combining a first image provided by the medical insurance business with semantic prompt text related to medical insurance product marketing. Based on the generation of these second images, which can be used for marketing medical insurance products, the convenience of marketing these products is improved. Of course, it is not limited to this, and no restrictions are set here.

[0125] Please refer to Figure 3, which is a schematic block diagram of a training apparatus for an image generation model provided in an embodiment of this application. This training apparatus can be configured in a server or computer device to execute the aforementioned training method for the image generation model.

[0126] As shown in Figure 3, the training device for the image generation model includes: a first acquisition module 110, a second acquisition module 120, a processing module 130, a first loss determination module 140, a second loss determination module 150, and a model training module 160.

[0127] The first acquisition module 110 is used to acquire a training dataset, which includes multiple first images and image annotation text corresponding to each first image.

[0128] The second acquisition module 120 is used to acquire semantic prompt text;

[0129] Processing module 130 is used to perform image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain a second image corresponding to the first image;

[0130] The first loss determination module 140 is used to determine the adversarial loss between the first image and the second image based on the first image and the second image;

[0131] The second loss determination module 150 is used to determine the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image;

[0132] The model training module 160 is used to adjust the model parameters of the image generation model based on the adversarial loss and the semantic control loss.

[0133] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and its modules and units can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0134] The method of this application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0135] For example, the above-described methods and apparatus can be implemented as a computer program that can run on a computer device.

[0136] Please refer to Figure 4, which is a schematic block diagram of a computer device provided in an embodiment of this application. This computer device can be a server or an electronic device.

[0137] As shown in Figure 4, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory may include storage media and internal memory.

[0138] The storage medium may store the operating system and computer programs. The computer programs include program instructions that, when executed, cause the processor to perform a training method for any image generation model.

[0139] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0140] Internal memory provides an environment for the execution of computer programs stored in the storage medium. When these computer programs are executed by the processor, the processor can perform any training method for image generation models.

[0141] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0142] It should be understood that a processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other convertible logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0143] In one embodiment, the processor is configured to execute a computer program and, when executing the computer program, perform the following steps:

[0144] Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image;

[0145] Obtain semantic prompt text;

[0146] Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image;

[0147] Based on the first image and the second image, determine the adversarial loss between the first image and the second image;

[0148] Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image;

[0149] The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

[0150] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific process of training the image generation model described above can be referred to the corresponding process in the aforementioned image generation model training method embodiments, and will not be repeated here.

[0151] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method implemented can be referred to in various embodiments of the training method for the image generation model of this application. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0152] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0153] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0154] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0155] It should also be understood that any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0156] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific implementations of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A training method for an image generation model, comprising: Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image; Obtain semantic prompt text; Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image; Based on the first image and the second image, determine the adversarial loss between the first image and the second image; Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image; The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

2. The training method according to claim 1, wherein, The step of determining the adversarial loss between the first image and the second image based on the first image and the second image includes: Based on the perceptual feature extraction network of the image generation model, perceptual feature extraction processing is performed on the first image and the second image respectively to obtain the first perceptual feature corresponding to the first image and the second perceptual feature corresponding to the second image. Determine the perceptual loss between the first perceptual feature and the second perceptual feature; Based on the residual network of the image generation model, semantic classification processing is performed on the first image and the second image respectively to obtain the first classification result corresponding to the first image and the second classification result corresponding to the second image; Determine the classification loss between the first classification result and the second classification result; The adversarial loss is determined based on the sum of the perceptual loss and the classification loss.

3. The training method according to claim 1, wherein, The step of determining the semantic control loss between the semantic cue text and the second image based on the semantic cue text and the second image includes: The text / image encoder based on the image generation model performs embedding feature extraction processing on the semantic prompt text and the second image respectively to obtain the first text embedding feature corresponding to the semantic prompt text and the second image embedding feature corresponding to the second image. The semantic control loss is determined based on the distance between the first text embedding feature and the second image embedding feature.

4. The training method according to any one of claims 1 to 3, wherein, The step of performing image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image includes: The text / image encoder based on the image generation model performs embedding feature extraction processing on the semantic prompt text, the image annotation text, and the first image respectively to obtain the first text embedding feature corresponding to the semantic prompt text, the second text embedding feature corresponding to the image annotation text, and the first image embedding feature corresponding to the first image. Based on the stable diffusion network of the image generation model and the semantic fine-tuning sub-network set in the stable diffusion network, the first image embedding feature is subjected to image semantic control processing according to the first text embedding feature and the second text embedding feature to obtain the second image corresponding to the first image.

5. The training method according to claim 4, wherein, The step of adjusting the model parameters of the image generation model based on the adversarial loss and the semantic control loss includes: Based on the adversarial loss and the semantic control loss, the preset network parameters of the semantic fine-tuning subnetwork are adjusted, while the preset network parameters of the subnetworks in the stable diffusion network other than the semantic fine-tuning subnetwork remain unchanged.

6. The training method according to any one of claims 1 to 3, wherein, The acquisition of the training dataset includes: Acquire multiple preset images; From multiple preset images, images that do not meet the preset image filtering criteria are removed to obtain the first image; The first image is subjected to image annotation processing to obtain the image annotation text corresponding to the first image.

7. The training method according to claim 6, wherein, After removing preset images that do not meet the preset image filtering criteria from multiple preset images to obtain the first image, the process further includes: The first image is subjected to image size transformation processing to obtain the first image under different image sizes; The step of performing image annotation generation processing on the first image to obtain the image annotation text corresponding to the first image includes: Image annotation generation is performed on the first image at different image sizes to obtain the corresponding image annotation text for the first image at different image sizes.

8. The training method according to any one of claims 1 to 3, wherein, The acquisition of semantic prompt text includes: Get the input text corresponding to the text input operation; If the input text meets the preset text filtering conditions, the input text is determined to be semantic prompt text; the preset text filtering conditions include at least one of the following: the input text is not garbled text and the input text is not pure punctuation marks.

9. The training method according to any one of claims 1 to 3, wherein, The step of performing image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image includes: Based on the semantic information of the first image and the image annotation text, determine whether the first image covers the object indicated by the semantic prompt text; If the first image does not cover the object indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that the object indicated by the semantic prompt text needs to be added to the first image, the object indicated by the semantic prompt text is added to the first image through the image generation model to obtain the second image corresponding to the first image; When the first image is covered by the object indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that the object indicated by the semantic prompt text needs to be deleted from the first image, the second image corresponding to the first image is obtained by deleting the object indicated by the semantic prompt text from the first image through the image generation model.

10. A training device for an image generation model, wherein, The training device includes: The first acquisition module is used to acquire a training dataset, which includes multiple first images and image annotation text corresponding to each first image. The second acquisition module is used to acquire semantic prompt text; The processing module is used to perform image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image; The first loss determination module is used to determine the adversarial loss between the first image and the second image based on the first image and the second image; The second loss determination module is used to determine the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image; The model training module is used to adjust the model parameters of the image generation model based on the adversarial loss and the semantic control loss.

11. A computer device, wherein, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, perform the following steps: Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image; Obtain semantic prompt text; Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image; Based on the first image and the second image, determine the adversarial loss between the first image and the second image; Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image; The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

12. A non-volatile computer-readable storage medium, wherein a computer program is stored on the non-volatile computer-readable storage medium, wherein, When the computer program is executed by the processor, it performs the following steps: Obtain a training dataset, which includes multiple first images and image annotation text corresponding to each first image; Obtain semantic prompt text; Based on the image generation model, the semantic prompt text, and the image annotation text, the first image is subjected to image semantic control processing to obtain the second image corresponding to the first image; Based on the first image and the second image, determine the adversarial loss between the first image and the second image; Based on the semantic cue text and the second image, determine the semantic control loss between the semantic cue text and the second image; The model parameters of the image generation model are adjusted based on the adversarial loss and the semantic control loss.

13. The non-volatile computer-readable storage medium according to claim 12, wherein, When the computer program is executed by the processor in the step of determining the adversarial loss between the first image and the second image based on the first image and the second image, the following steps are implemented: Based on the perceptual feature extraction network of the image generation model, perceptual feature extraction processing is performed on the first image and the second image respectively to obtain the first perceptual feature corresponding to the first image and the second perceptual feature corresponding to the second image. Determine the perceptual loss between the first perceptual feature and the second perceptual feature; Based on the residual network of the image generation model, semantic classification processing is performed on the first image and the second image respectively to obtain the first classification result corresponding to the first image and the second classification result corresponding to the second image; Determine the classification loss between the first classification result and the second classification result; The adversarial loss is determined based on the sum of the perceptual loss and the classification loss.

14. The non-volatile computer-readable storage medium according to claim 12, wherein, When the computer program is executed by the processor in the step of determining the semantic control loss between the semantic prompt text and the second image based on the semantic prompt text and the second image, the following steps are implemented: The text / image encoder based on the image generation model performs embedding feature extraction processing on the semantic prompt text and the second image respectively to obtain the first text embedding feature corresponding to the semantic prompt text and the second image embedding feature corresponding to the second image. The semantic control loss is determined based on the distance between the first text embedding feature and the second image embedding feature.

15. The non-volatile computer-readable storage medium according to any one of claims 12 to 14, wherein, When the computer program is executed by the processor to perform image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image, the following steps are implemented: The text / image encoder based on the image generation model performs embedding feature extraction processing on the semantic prompt text, the image annotation text, and the first image respectively to obtain the first text embedding feature corresponding to the semantic prompt text, the second text embedding feature corresponding to the image annotation text, and the first image embedding feature corresponding to the first image. Based on the stable diffusion network of the image generation model and the semantic fine-tuning sub-network set in the stable diffusion network, the first image embedding feature is subjected to image semantic control processing according to the first text embedding feature and the second text embedding feature to obtain the second image corresponding to the first image.

16. The non-volatile computer-readable storage medium according to claim 15, wherein, When the computer program is executed by the processor to adjust the model parameters of the image generation model based on the adversarial loss and the semantic control loss, the following steps are implemented: Based on the adversarial loss and the semantic control loss, the preset network parameters of the semantic fine-tuning subnetwork are adjusted, while the preset network parameters of the subnetworks in the stable diffusion network other than the semantic fine-tuning subnetwork remain unchanged.

17. The non-volatile computer-readable storage medium according to any one of claims 12 to 14, wherein, When the computer program is executed by the processor in the step of acquiring the training dataset, it implements the following steps: Acquire multiple preset images; From multiple preset images, images that do not meet the preset image filtering criteria are removed to obtain the first image; The first image is subjected to image annotation processing to obtain the image annotation text corresponding to the first image.

18. The non-volatile computer-readable storage medium according to claim 17, wherein, After the computer program is executed by the processor in the step of removing preset images from multiple preset images that do not meet the preset image filtering conditions to obtain the first image, it performs the following steps: The first image is subjected to image size transformation processing to obtain the first image under different image sizes; The step of performing image annotation generation processing on the first image to obtain the image annotation text corresponding to the first image includes: Image annotation generation is performed on the first image at different image sizes to obtain the corresponding image annotation text for the first image at different image sizes.

19. The non-volatile computer-readable storage medium according to any one of claims 12 to 14, wherein, When the computer program is executed by the processor in the step of obtaining semantic prompt text, it implements the following steps: Get the input text corresponding to the text input operation; If the input text meets the preset text filtering conditions, the input text is determined to be semantic prompt text; the preset text filtering conditions include at least one of the following: the input text is not garbled text and the input text is not pure punctuation marks.

20. The non-volatile computer-readable storage medium according to any one of claims 12 to 14, wherein, When the computer program is executed by the processor to perform image semantic control processing on the first image based on the image generation model, the semantic prompt text, and the image annotation text to obtain the second image corresponding to the first image, the following steps are implemented: Based on the semantic information of the first image and the image annotation text, determine whether the first image covers the object indicated by the semantic prompt text; If the first image does not cover the object indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that the object indicated by the semantic prompt text needs to be added to the first image, the object indicated by the semantic prompt text is added to the first image through the image generation model to obtain the second image corresponding to the first image; When the first image is covered by the object indicated by the semantic prompt text, and the semantic information of the semantic prompt text indicates that the object indicated by the semantic prompt text needs to be deleted from the first image, the second image corresponding to the first image is obtained by deleting the object indicated by the semantic prompt text from the first image through the image generation model.