Image generation model training method and device, equipment and storage medium
By introducing semantic hint text and adversarial loss mechanisms into the image generation model and adjusting model parameters, the problem of poor semantic control accuracy of image generation model is solved, and higher semantic control accuracy and training effect are achieved.
Patent Information
- Application Number
- CN202510066870.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
During the image generation process, the image generation model is prone to problems that the objects contained in the generated image do not match the objects indicated by the text, resulting in poor semantic control accuracy.
By obtaining the training data set and semantic prompt text, and performing semantic control processing on the image based on the image generation model, the adversarial loss and semantic control loss are determined, and the model parameters are adjusted to improve the accuracy of semantic control.
Improve the semantic control accuracy of the image generation model during the image generation process, ensuring that the generated image matches the object indicated by the text, thereby improving the training effect.
Smart Images

Figure CN119992252A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence model training technology, and in particular to a training method, device, equipment and storage medium for an image generation model. Background Art
[0002] At present, in the related art, images can be generated by combining image generation models, such as the Stable Diffusion (SD) model, with corresponding text. However, in the process of generating images through image generation models, it is easy for the objects contained in the generated images to be inconsistent with the objects indicated by the text, resulting in poor semantic control accuracy of the image generation model in the process of generating images. Summary of the invention
[0003] The main purpose of this application is to provide a training method, device, equipment and storage medium for an image generation model, aiming to improve the semantic control accuracy of the image generation model during the image generation process, so as to improve the training effect of the image generation model.
[0004] In a first aspect, the present application provides a method for training an image generation model, comprising:
[0005] Acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images;
[0006] Get semantic hint text;
[0007] Based on the image generation model, the semantic prompt text and the image annotation text, performing image semantic control processing on the first image to obtain a second image corresponding to the first image;
[0008] Determine, according to the first image and the second image, an adversarial loss between the first image and the second image;
[0009] Determining a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image;
[0010] According to the adversarial loss and the semantic control loss, model parameters of the image generation model are adjusted.
[0011] In a second aspect, the present application further provides a training device for an image generation model, the training device comprising:
[0012] A first acquisition module is used to acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images;
[0013] The second acquisition module is used to acquire the semantic prompt text;
[0014] A processing module, configured to perform image semantic control processing on the first image based on an image generation model, the semantic prompt text, and the image annotation text, to obtain a second image corresponding to the first image;
[0015] A first loss determination module, configured to determine an adversarial loss between the first image and the second image according to the first image and the second image;
[0016] A second loss determination module, configured to determine a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image;
[0017] A model training module is used to adjust the model parameters of the image generation model according to the adversarial loss and the semantic control loss.
[0018] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor;
[0019] The memory is used to store computer programs;
[0020] The processor is used to execute the computer program and implement the steps of the training method of the image generation model as described above when executing the computer program.
[0021] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the image generation model as described above are implemented.
[0022] The present application provides a training method, apparatus, device and storage medium for an image generation model, the training method comprising: obtaining a training data set, the training data set comprising a plurality of first images and image annotation texts corresponding to each of the first images; obtaining a semantic prompt text; performing image semantic control processing on the first image based on the image generation model, the semantic prompt text and the image annotation text to obtain a second image corresponding to the first image; determining an adversarial loss between the first image and the second image according to the first image and the second image; determining a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image; and adjusting model parameters of the image generation model according to the adversarial loss and the semantic control loss.
[0023] For example, in the process of training the image generation model, the model parameters of the image generation model can be adjusted according to the adversarial loss and the semantic control loss. Since the adversarial loss is determined according to the first image and the second image, the image generation model can evaluate whether the second image is similar to the first image according to the adversarial loss. Correspondingly, since the semantic control loss is determined according to the semantic prompt text and the second image, the image generation model can evaluate whether the object contained in the second image is consistent with the object indicated by the semantic prompt text according to the semantic control loss. Based on this, when the model parameters of the image generation model are adjusted according to the adversarial loss and the semantic control loss, the image generation model can perform image semantic control processing on the first image according to the semantic prompt text to obtain a second image that is similar to the first image and contains objects that are consistent with the objects indicated by the semantic prompt text, which is conducive to improving the semantic control accuracy of the image generation model in the process of image generation, so as to improve the training effect of the image generation model.
[0024] The trained image generation model obtained by the training method of the image generation model can be applied to the fields of financial technology, medical health, etc. In some embodiments, for the auto insurance business in the field of financial technology, in response to the marketing needs of the auto insurance business for auto insurance products, for example, a large number of auto insurance product marketing images related to cars need to be generated, then the trained image generation model can be combined with the first image provided by the auto insurance business and the semantic prompt text related to the marketing of auto insurance products to generate a large number of second images similar to the first image and containing objects that match the objects indicated by the semantic prompt text. Based on the generation of the second image, the second image can be used for the marketing of auto insurance products, which is conducive to improving the convenience of marketing of auto insurance products. In other embodiments, for the medical insurance business in the field of medical health, in response to the marketing needs of the medical insurance business for medical insurance products, for example, a large number of medical insurance product marketing images related to medical care need to be generated, then the trained image generation model can be combined with the first image provided by the medical insurance business and the semantic prompt text related to the marketing of medical insurance products to generate a large number of second images similar to the first image and containing objects that match the objects indicated by the semantic prompt text. Based on the generation of the second image, the second image can be used for the marketing of medical insurance products, which is conducive to improving the convenience of marketing of medical insurance products. Of course, it is not limited to this and no limitation is made here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 It is a flowchart of a training method for an image generation model provided in an embodiment of the present application;
[0027] Figure 2 is a flowchart of an image generation method involved in an embodiment of the present application;
[0028] Figure 3 is a schematic block diagram of a training device for an image generation model provided in an embodiment of the present application;
[0029] Figure 4 It is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0031] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0032] The embodiments of the present application provide a training method, apparatus, device and storage medium for an image generation model. Among them, the training method of the image generation model can be applied to a computer device, which can be a tablet computer, a laptop computer, a desktop computer and other devices. It can also be applied to a server, which can be a separate server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0033] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0034] See also Figure 1 , Figure 1It is a flow chart of a training method for an image generation model provided by an embodiment of the present application. It should be noted that the training method for the image generation model provided by the embodiment of the present application can be used for a computer device, and of course it can also be used for a server. For example, the server can obtain a training data set from a computer device, and process the training data set accordingly according to the training method of the image generation model to obtain a second image corresponding to the first image, and then determine the adversarial loss between the first image and the second image and determine the semantic control loss between the semantic prompt text and the second image, so as to adjust the model parameters of the image generation model according to the adversarial loss and the semantic control loss. Exemplarily, the server can, for example, send a trained image generation model obtained according to the training method of the image generation model to a computer device, so that the computer device generates a corresponding second image based on the image generation model; of course, it is not limited to this, and the trained image generation model can be applied to different fields such as the financial technology field, the medical health field, etc. to generate corresponding images, which is not limited here.
[0035] In specific implementation, the computer device includes but is not limited to: any one of a tablet computer, a laptop computer, and a desktop computer; the server can be a single server or a server cluster, or a cloud server that provides cloud computing services.
[0036] like Figure 1 As shown, the training method of the image generation model includes steps S101 to S106.
[0037] Step S101: Acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each first image.
[0038] For example, a computer device may acquire a plurality of first images to construct a training data set based on the plurality of first images. The training data set may be used by the computer device to train the image generation model. When the training data set includes a first image, the computer device may input the first image into the image generation model, and the image generation model may use the first image as a basis for subsequently generating a second image corresponding to the first image. The objects included in different first images may be the same or different, which is not limited here. Among them, the objects in the first image may include text, objects, etc., which is not limited here. Texts include, for example, artistic characters, text in preset fonts, letters, numbers, etc., which are not limited here. Objects include, for example, cars, electrical appliances, daily necessities, etc., which are not limited here.
[0039] In an exemplary embodiment, when training an image generation model, a computer device can obtain images related to different fields such as the financial technology field, the medical health field, etc. based on the consideration that the trained image generation model can be subsequently applied to different fields such as the financial technology field, the medical health field, etc., as multiple first images included in the training data set. Images related to the financial technology field include, for example, vehicle images involved in auto insurance business, auto insurance scene images, claims images related to claims information involved in insurance claims business, etc., which are not limited here. Images related to the medical health field include, for example, images related to different types of diseases, images related to different types of disease symptoms, etc., which are not limited here. Of course, when the image generation model is subsequently trained, it can also be based on images that are not related to different fields such as the financial technology field, the medical health field, etc. to train the image generation model on how to generate images related to different fields such as the financial technology field, the medical health field, etc. according to the instructions of the semantic prompt text, which is not limited here.
[0040] When a plurality of first images are acquired, the computer device may determine the image annotation text corresponding to each of the first images, so as to construct a training data set in combination with the first image and the image annotation text corresponding to the first image. The image annotation text may be used to indicate the object included in the first image. For example, the computer device may perform image annotation processing on the first image through a preset image annotation tool to obtain the image annotation text corresponding to the first image. The image annotation tool may include a multimodal visual-text large language model, such as a BLIP2 model, etc., which is not limited here. For example, the image annotation tool may identify objects such as text, objects, etc. in the first image to obtain a corresponding recognition result. Accordingly, the image annotation tool may set the recognition result as the image annotation text corresponding to the first image in the form of text.
[0041] When the training data set is obtained, the computer device can subsequently train how to improve the semantic control accuracy of the image generation model during the image generation process based on the multiple first images included in the training data set and the image annotation texts corresponding to each first image.
[0042] In some embodiments, a plurality of preset images are obtained; preset images that do not meet the preset image screening conditions are eliminated from the plurality of preset images to obtain a first image; and image annotation processing is performed on the first image to obtain image annotation text corresponding to the first image.
[0043] In the process of training the image generation model to improve the semantic control accuracy of the image generation model in the process of image generation, the computer device can obtain as many first images as possible for use by the image generation model. Based on this, the computer device can obtain multiple preset images to filter and / or process the first image from the multiple preset images. For example, the computer device can obtain multiple preset images based on public data, such as open source data sets, websites, etc. Correspondingly, the computer device can determine whether there are preset images that do not meet the preset image screening conditions among the multiple preset images, so as to eliminate the preset images that do not meet the preset image screening conditions and obtain the first image. The preset image screening conditions may include at least one of the presence of a watermark in the preset image and the copyright of the preset image. Of course, the preset image screening conditions are not limited to this. The preset image screening conditions can be pre-set or set by the user, which is not limited here.
[0044] When the first image is acquired, the computer device may determine the image annotation text corresponding to the first image, for example, by performing image annotation processing on the first image through a preset image annotation tool to obtain the image annotation text corresponding to the first image.
[0045] In this way, the computer device can filter out the first image from multiple preset images according to the preset image screening conditions, which is conducive to improving the convenience of determining the first image. Accordingly, the first image can be used for the computer device to perform image annotation processing to determine the image annotation text corresponding to the first image. The image annotation text can be used to indicate the object included in the first image, which can be used for subsequent training of the image generation model.
[0046] Exemplarily, after acquiring the first image, the computer device may process the first image accordingly to utilize a limited number of first images to acquire a larger number of first images for subsequent training of an image generation model.
[0047] For example, the first image is subjected to image size conversion processing to obtain the first image in a different image size.
[0048] The computer device may perform image size conversion processing on the first image. For example, the computer device may convert the first image into first images of different image sizes. Different image sizes include, for example, 1024 pixels × 1024 pixels, 768 pixels × 768 pixels, 1024 pixels × 768 pixels, etc., which are not limited here. Different image sizes may be pre-set or may be set by the user, which are not limited here. In the case where the first image can be processed by image size conversion to obtain first images of different image sizes, the number of first images of different image sizes may be more than the number of original first images, and the first images of different image sizes may all be used for subsequent training of the image generation model, which is beneficial to increase the amount of training data of the image generation model, so as to improve the training effect of the image generation model in the future.
[0049] For example, image annotation generation processing is performed on the first images at different image sizes respectively to obtain image annotation texts corresponding to the first images at different image sizes.
[0050] When the first images at different image sizes are obtained, the computer device may determine the image annotation texts corresponding to the first images at different image sizes. For example, by using a preset image annotation tool, the first images at different image sizes are respectively annotated to obtain the image annotation texts corresponding to the first images at different image sizes. The image annotation texts may be used to indicate the objects included in the first images at different image sizes for subsequent training of the image generation model.
[0051] Step S102: Obtain semantic prompt text.
[0052] For example, a computer device may obtain semantic prompt text during the process of training an image generation model. For example, the semantic prompt text may be input by a relevant person, and when the computer device detects a text input operation by the relevant person, it may obtain the input text corresponding to the text input operation. If the input text does not meet the preset text filtering conditions, the input text may be eliminated. If the input text meets the preset text filtering conditions, the input text may be determined as semantic prompt text. The preset text filtering conditions may include that the input text is not garbled, the input text is not pure punctuation, etc., which are not limited here. The preset text filtering conditions may be pre-set or may be set by the user, which are not limited here.
[0053] The semantic hint text may be used by the image generation model to determine the semantic control requirements when performing image semantic control processing on the first image. The semantic hint text may be used to indicate a change to an object included in the first image. The change to an object included in the first image may include, for example, adding a new object to the first image and / or deleting an old object from the first image.
[0054] Take the auto insurance business in the field of financial technology as an example. In response to the marketing needs of the auto insurance business for auto insurance products, the auto insurance business requires that the image generation model can provide a large number of auto insurance product marketing images, and the computer device can determine the semantic prompt text based on the objects related to the auto insurance products. Objects related to auto insurance products include, for example, auto insurance customer categories, types of cars covered by auto insurance, auto insurance types, auto insurance scenarios, etc., and the computer device can determine the auto insurance customer categories, types of cars covered by auto insurance, auto insurance types, auto insurance scenarios, etc. as semantic prompt texts to input into the image generation model.
[0055] Take the medical insurance business in the medical and health field as an example. In response to the marketing needs of the medical insurance business for medical insurance products, the medical insurance business requires that the image generation model can provide a large number of marketing images of medical insurance products. The computer device can determine the semantic prompt text based on the objects related to the medical insurance products. The objects related to the medical insurance products include, for example, the categories of medical insurance customers, the types of diseases covered by medical insurance, the types of medical insurance, etc. The computer device can determine the categories of medical insurance customers, the types of diseases covered by medical insurance, the types of medical insurance, etc. as semantic prompt texts to input into the image generation model.
[0056] The semantic hint text can be combined with the semantic hint text by the image generation model to perform image semantic control processing on the first image to obtain a second image corresponding to the first image.
[0057] In this way, when the semantic hint text is obtained, the image generation model can perform image semantic control processing on the first image based on the semantic hint text, so as to subsequently determine the second image corresponding to the first image, which is conducive to improving the convenience of subsequent image semantic control processing of the first image.
[0058] Step S103: performing image semantic control processing on the first image based on the image generation model, the semantic prompt text and the image annotation text to obtain a second image corresponding to the first image.
[0059] When the computer device obtains the first image and the semantic hint text, it can input the first image and the semantic hint text into the image generation model. Since the first image can carry the image annotation text corresponding to the first image, the image annotation text corresponding to the first image can also be input into the image generation model accordingly. When the image generation model receives the first image, the semantic hint text and the image annotation text corresponding to the first image, it can combine the semantic hint text and the image annotation text to perform image semantic control processing on the first image to obtain a second image corresponding to the first image.
[0060] In some embodiments, the image generation model can perform image semantic control processing on the first image based on the semantic information of the semantic prompt text. The image annotation text can provide additional semantic information for the first image. The image generation model can determine the semantic information of the semantic prompt text and the image annotation text. For example, the image generation model can use natural language processing techniques, such as word embedding, semantic analysis, etc., to convert the semantic prompt text and the image annotation text into representations that can be understood by the image generation model, thereby obtaining the semantic information of the semantic prompt text and the image annotation text.
[0061] Accordingly, in the process of performing image semantic control processing on the first image by the image generation model, the image generation model can determine whether the first image is covered with the object indicated by the semantic hint text based on the semantic information of the first image and the image annotation text corresponding to the first image, so as to further change the object in the first image in combination with the semantic information of the semantic hint text. For example, when it is determined that the first image is not covered with the object 1 indicated by the semantic hint text in combination with the semantic information of the first image and the image annotation text, and the semantic information of the semantic hint text indicates that the object 1 needs to be added to the first image, the image generation model can add the object 1 to the first image to determine the second image corresponding to the first image. For another example, when it is determined that the first image is covered with the object 2 indicated by the semantic hint text in combination with the semantic information of the first image and the image annotation text, and the semantic information of the semantic hint text indicates that the object 2 needs to be deleted from the first image, the image generation model can delete the object 2 in the first image to determine the second image corresponding to the first image. Of course, it is not limited to this, and no limitation is made here.
[0062] In this way, based on the image generation model, semantic prompt text and image annotation text, when the first image is processed by image semantic control and the second image corresponding to the first image is obtained, the image generation model can be trained on how to improve the accuracy of semantic control in the image generation process, so that the second image is similar to the first image, and the object contained in the second image is consistent with the object indicated by the semantic prompt text, thereby improving the training effect of the image generation model.
[0063] Exemplarily, the image generation model may include a text / image encoder and a stable diffusion model provided with a semantic fine-tuning subnetwork. In an exemplary embodiment, the text / image encoder may adopt open-clip-g. The stable diffusion model may adopt SDXL1.0. The semantic fine-tuning subnetwork may include a LoRA subnetwork. The LoRA subnetwork may be arranged in the attention structure of the stable diffusion network. Among them, the attention structure of the stable diffusion subnetwork may be arranged in a U-Net architecture. The U-Net architecture, for example, includes a convolutional layer, an attention layer, an upsampling layer, and a downsampling layer, then the LoRA subnetwork may be arranged in one of the attention layers of the U-Net architecture. Of course, it is not limited to this and is not limited here.
[0064] In some embodiments, a text / image encoder based on an image generation model performs embedding feature extraction processing on the semantic prompt text, the image annotation text, and the first image, respectively, to obtain a first text embedding feature corresponding to the semantic prompt text, a second text embedding feature corresponding to the image annotation text, and a first image embedding feature corresponding to the first image; a stable diffusion network based on the image generation model and a semantic fine-tuning subnetwork set in the stable diffusion network perform image semantic control processing on the first image embedding feature according to the first text embedding feature and the second text embedding feature, to obtain a second image corresponding to the first image.
[0065] For example, the image generation model can input the acquired semantic hint text, image annotation text and the first image into a text / image encoder. The text / image encoder can have the ability to perform embedding feature extraction processing on text and image, and the text / image encoder can respectively perform embedding feature extraction processing on the semantic hint text, image annotation text and the first image, to obtain the first text embedding feature corresponding to the semantic hint text, the second text embedding feature corresponding to the image annotation text and the first image embedding feature corresponding to the first image. The first text embedding feature and the second text embedding feature can be used by the image generation model to determine whether the first image is covered by the object indicated by the semantic hint text, so as to provide subsequent instructions for image semantic control processing of the first image. The first image embedding feature can be used by the image generation model to perform corresponding image semantic control processing, and then determine the second image corresponding to the first image.
[0066] When the first text embedding feature, the second text embedding feature, and the first image embedding feature are determined, the image generation model may input the first text embedding feature, the second text embedding feature, and the first image embedding feature into a stable diffusion network. The stable diffusion network may have the ability to generate images based on text, and the semantic fine-tuning subnetwork provided in the stable diffusion network may have the ability to fine-tune the stable diffusion network. Then, the stable diffusion network provided with the semantic fine-tuning subnetwork may be provided for the image generation model to perform image semantic control processing on the first image embedding feature based on the first text embedding feature and the second text embedding feature, so as to obtain the second image corresponding to the first image based on the fine-tuned first image embedding feature. For example, the stable diffusion network provided with the semantic fine-tuning subnetwork may perform image semantic control processing on the first image embedding feature based on the semantic information of the first text embedding feature and the second text embedding feature, so as to change the objects included in the first image in response to the instruction of the semantic prompt text, so that the second image matches the semantic information of the semantic prompt text.
[0067] Take the auto insurance business in the field of financial technology as an example. The first image includes, for example, at least one of a vehicle image and a non-vehicle image, and the image annotation text corresponding to the first image can be used to indicate that the object in the vehicle image includes a vehicle, and the object in the non-vehicle image does not include a vehicle. The objects indicated by the semantic prompt text include, for example, the category of auto insurance customers, the type of car covered by auto insurance, the type of auto insurance, and the auto insurance scenario, and the semantic information of the semantic prompt text is used to indicate the object indicated by the semantic prompt text added in the first image. The image generation model can use a text / image editor to perform embedded feature extraction processing on the image annotation text and the semantic prompt text corresponding to the vehicle image, the non-vehicle image, the vehicle image, and the non-vehicle image, respectively, to determine the first image embedding features corresponding to the vehicle image and the non-vehicle image, the first text embedding features corresponding to the semantic prompt text, and the second text embedding features corresponding to the image annotation text. The image generation model can judge whether the objects in the vehicle image and the non-vehicle image include the category of auto insurance customers, the type of cars covered by auto insurance, the type of auto insurance, and the auto insurance scene, respectively, based on the first text embedding feature and the second text embedding feature, so that when the vehicle image and / or the non-vehicle image are not covered with the corresponding object, the image semantic control processing is performed on the vehicle image and the non-vehicle image respectively by combining the first text embedding feature and the second text embedding feature through a stable diffusion network provided with a semantic fine-tuning subnetwork, thereby obtaining the second image corresponding to the vehicle image and the non-vehicle image. Among them, the objects in the second images corresponding to the vehicle image and the non-vehicle image can be added with at least one of the category of auto insurance customers, the type of cars covered by auto insurance, the type of auto insurance, and the auto insurance scene according to the indication of the semantic information of the semantic hint text. Of course, it is not limited to this and is not limited here.
[0068] Take the medical insurance business in the medical and health field as an example. The first image includes, for example, at least one of a disease image and a non-disease image, and the image annotation text corresponding to the first image can be used to indicate that the objects in the disease image include the corresponding disease type, and the objects in the non-disease image do not include any disease type. The objects indicated by the semantic prompt text include, for example, the disease types covered by medical insurance, the categories of medical insurance customers, and the types of medical insurance, and the semantic information of the semantic prompt text is used to indicate the objects indicated by the semantic prompt text added in the first image. The image generation model can use a text / image editor to perform embedded feature extraction processing on the image annotation text and the semantic prompt text corresponding to the disease image, non-disease image, disease image, and non-disease image, respectively, to determine the first image embedding features corresponding to the disease image and the non-disease image, the first text embedding features corresponding to the semantic prompt text, and the second text embedding features corresponding to the image annotation text. The image generation model can judge whether the objects in the disease image and the non-disease image include the types of diseases covered by medical insurance, the categories of medical insurance customers, and the types of medical insurance based on the first text embedding features and the second text embedding features, so that when the disease image and / or the non-disease image are not covered with the corresponding objects, the image semantic control processing is performed on the disease image and the non-disease image respectively by combining the first text embedding features and the second text embedding features through a stable diffusion network provided with a semantic fine-tuning subnetwork, thereby obtaining the second images corresponding to the disease image and the non-disease image. Among them, the objects in the second images corresponding to the disease image and the non-disease image can be added with at least one of the types of diseases covered by medical insurance, the categories of medical insurance customers, and the types of medical insurance according to the instructions of the semantic information of the semantic hint text. Of course, it is not limited to this and is not limited here.
[0069] In this way, the image generation model can determine the first text embedding feature corresponding to the semantic prompt text, the second text embedding feature corresponding to the image annotation text, and the first image embedding feature corresponding to the first image through a text / image encoder, and perform image semantic control processing on the first image embedding feature according to the first text embedding feature and the second text embedding feature through a stable diffusion network provided with a semantic fine-tuning subnetwork to obtain a second image corresponding to the first image. The image generation model can then determine whether the object indicated by the semantic prompt text is consistent with the object indicated by the image annotation text based on the first text embedding feature and the second text embedding feature, so that when there is a difference, image semantic control processing is performed on the first image according to the object indicated by the semantic prompt text, thereby making the second image corresponding to the first image match the semantic information of the semantic prompt text. This is beneficial to training the image generation model on how to improve the accuracy of semantic control of the image during the image generation process, thereby improving the training effect of the image generation model.
[0070] Step S104: determine the adversarial loss between the first image and the second image according to the first image and the second image.
[0071] Exemplarily, when training an image generation model on how to improve the accuracy of semantic control of an image during image generation, a computer device may introduce a generative adversarial network (GAN) adversarial mechanism into the image generation model to combine the adversarial mechanism of the GAN network to determine the adversarial loss between the first image and the second image. For example, the GAN network needs to determine the loss of the generator and the loss of the discriminator when determining the adversarial loss. The goal of the generator is to generate realistic fake images to deceive the discriminator, and the loss of the generator may be associated with the probability that the discriminator judges the fake image as a real image. The goal of the discriminator is to accurately distinguish between real images and fake images, and the loss of the discriminator may be associated with the probability that the discriminator judges the real image as a real image and the probability that the fake image is judged as a fake image. Based on this, the image generation model may set the first image as a real image involved in the GAN network, and set the second image as a fake image involved in the GAN network, so that the image generation model can determine the adversarial loss between the first image and the second image. For example, the image generation model can determine the adversarial loss by evaluating the probability of judging the first image to be the second image by judging whether the first image is similar to the second image.
[0072] In the process of judging whether the first image is similar to the second image, the image generation model can evaluate the probability of judging the first image to be the second image by judging the image similarity between the first image and the second image, as well as the semantic similarity between the first image and the second image, thereby determining the adversarial loss between the first image and the second image.
[0073] In some embodiments, based on a perceptual feature extraction network of an image generation model, perceptual feature extraction processing is performed on the first image and the second image, respectively, to obtain a first perceptual feature corresponding to the first image and a second perceptual feature corresponding to the second image; the perceptual loss between the first perceptual feature and the second perceptual feature is determined; based on a residual network of the image generation model, semantic classification processing is performed on the first image and the second image, respectively, to obtain a first classification result corresponding to the first image and a second classification result corresponding to the second image; the classification loss between the first classification result and the second classification result is determined; and the adversarial loss is determined based on the sum of the perceptual loss and the classification loss.
[0074] The image generation model can be provided with a perceptual feature extraction network. The perceptual feature extraction network can be used to perform perceptual feature extraction processing on an image to obtain perceptual features corresponding to the image. Perceptual features can be used to indicate features in an image that can be perceived and recognized by the human visual system. Perceptual features can be related to the visual content of the image, and the perceptual features include, for example, color, texture, shape, edge, and higher-level semantic information, such as objects, scenes, etc., which are not limited here. Perceptual feature extraction networks include, for example, deep learning networks such as VGG (Visual Geometry Group) networks, which are not limited here.
[0075] For example, the image generation model can input the first image and the second image into the perceptual feature extraction network, so that the perceptual feature extraction network performs perceptual feature extraction processing on the first image and the second image respectively to obtain the first perceptual feature corresponding to the first image and the second perceptual feature corresponding to the second image.
[0076] Accordingly, the first perceptual feature and the second perceptual feature may be used by the image generation model to determine the image similarity between the first image and the second image. For example, the image generation model may determine the difference between the first perceptual feature and the second perceptual feature to determine the perceptual loss between the first perceptual feature and the second perceptual feature based on the difference between the first perceptual feature and the second perceptual feature. The perceptual loss may be positively correlated with the image similarity between the first image and the second image. The smaller the perceptual loss, the higher the image similarity between the first image and the second image, and the larger the perceptual loss, the lower the image similarity between the first image and the second image.
[0077] The image generation model may be provided with a residual network (ResNet). The residual network may be used to perform semantic classification processing on the image to obtain a classification result corresponding to the image. The classification result may be used to indicate a semantic classification result corresponding to the semantic information of the image. The classification result may be related to the semantic information of the image, and the classification result may include, for example, an object classification result, a scene classification result, etc., which is not limited here.
[0078] For example, the image generation model can input the first image and the second image into the residual network, so that the residual network performs semantic classification processing on the first image and the second image respectively to obtain a first classification result corresponding to the first image and a second classification result corresponding to the second image.
[0079] Accordingly, the first classification result and the second classification result may be used by the image generation model to determine the semantic similarity between the first image and the second image. For example, the image generation model may determine the difference between the first classification result and the second classification result to determine whether the first classification result and the second classification result belong to the same category, and then determine the classification loss between the first classification result and the second classification result. The classification loss may be minimized when the first classification result and the second classification result belong to the same category. The smaller the classification loss, the higher the semantic similarity between the first image and the second image, and the larger the classification loss, the lower the semantic similarity between the first image and the second image.
[0080] When determining the perceptual loss between the first perceptual feature and the second perceptual feature, and the classification loss between the first classification result and the second classification result, the image generation model may determine the adversarial loss based on the sum of the perceptual loss and the classification loss. For example, when training the image generation model, the model parameters of the image generation model may be adjusted based on a strategy of minimizing the adversarial loss, i.e., minimizing the sum of the perceptual loss and the classification loss, so that the second image generated by the image generation model is similar to the first image input to the image generation model.
[0081] In this way, the image generation model can determine the adversarial loss of the image generation model based on the first image and the second image. The adversarial loss can be used for subsequent training of the image generation model to improve the similarity between the second image generated by the image generation model and the first image input to the image generation model, which is conducive to subsequent improvement of the training effect of the image generation model.
[0082] Step S105: determining the semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image.
[0083] Exemplarily, when training an image generation model on how to improve the accuracy of semantic control of an image during the image generation process, a computer device can determine, based on the semantic prompt text and the second image, whether the semantic information of the second image is consistent with the semantic information of the semantic prompt text, so as to determine the accuracy of image semantic control when the image generation model performs image semantic control processing on the first image in response to the instructions of the semantic prompt text.
[0084] In some embodiments, a text / image encoder based on an image generation model performs embedding feature extraction processing on the semantic hint text and the second image, respectively, to obtain a first text embedding feature corresponding to the semantic hint text and a second image embedding feature corresponding to the second image; and the semantic control loss is determined based on the distance between the first text embedding feature and the second image embedding feature.
[0085] The image generation model may be provided with a text / image encoder. The text / image encoder may be used to perform embedding feature extraction processing on at least one of the text and the image to obtain embedding features corresponding to at least one of the text and the image. The embedding features may be used to indicate semantic information of at least one of the text and the image.
[0086] For example, the image generation model can input the semantic prompt text and the second image into the text / image encoder, so that the text / image encoder performs embedding feature extraction processing on the semantic prompt text and the second image respectively to obtain the first text embedding feature corresponding to the semantic prompt text and the second image embedding feature corresponding to the second image.
[0087] Accordingly, the first text embedding feature and the second image embedding feature can be used by the image generation model to determine whether the semantic information of the second image is consistent with the semantic information of the semantic prompt text. For example, the higher the degree of match between the object in the second image and the object indicated by the semantic prompt text, the higher the possibility that the semantic information of the second image is consistent with the semantic information of the semantic prompt text. Based on this, the image generation model can determine the distance between the first text embedding feature and the second image embedding feature to evaluate the degree of match between the object in the second image and the object indicated by the semantic prompt text, thereby obtaining the semantic control loss. For example, the image generation model can determine the semantic control loss by calculating a measurable distance such as cosine similarity between the first text embedding feature and the second image embedding feature, which is not limited here.
[0088] When the semantic control loss is determined based on the distance between the first text embedding feature and the second image embedding feature, the semantic control accuracy of the image generation model in the process of generating the image can be evaluated based on the semantic control loss. The larger the semantic control loss, the worse the semantic control accuracy of the image generation model. The smaller the semantic control loss, the better the semantic control accuracy of the image generation model. For example, when training the image generation model, the model parameters of the image generation model can be adjusted based on the strategy of minimizing the semantic control loss, that is, minimizing the distance between the first text embedding feature and the second image embedding feature, so that the object indicated by the second image generated by the image generation model is consistent with the object indicated by the semantic prompt text.
[0089] In this way, the image generation model can determine the semantic control loss of the image generation model according to the semantic hint text and the second image. The semantic control loss can be used for subsequent training of the image generation model to improve the matching degree between the object indicated by the second image generated by the image generation model and the object indicated by the semantic hint text, which is conducive to the subsequent improvement of the semantic control accuracy of the image generation model in the image generation process, so as to improve the training effect of the image generation model.
[0090] Step S106: adjusting the model parameters of the image generation model according to the adversarial loss and the semantic control loss.
[0091] When the adversarial loss and semantic control loss of the image generation model are determined, the model parameters of the image generation model can be adjusted according to the object loss and the semantic control loss. For example, the image generation model can adjust the model parameters of the image generation model according to the strategy of minimizing the adversarial loss and the semantic control loss to obtain a trained image generation model. Accordingly, when the trained image generation model is subsequently used to generate a second image corresponding to the corresponding first image, it is beneficial to improve the semantic control accuracy of the image generation model in the process of generating the second image, thereby improving the image generation accuracy and image generation effect of the second image.
[0092] In some embodiments, the preset network parameters of the semantic fine-tuning subnetwork are adjusted according to the adversarial loss and the semantic control loss, and the preset network parameters of the subnetworks other than the semantic fine-tuning subnetwork in the stable diffusion network are kept unchanged.
[0093] For example, when training an image generation model to improve the semantic control accuracy of the image generation model, the preset network parameters of the subnetworks other than the semantic fine-tuning subnetwork in the stable diffusion network can be frozen, and only the preset network parameters of the semantic fine-tuning subnetwork can be adjusted and updated. When training the image generation model, the preset network parameters of the semantic fine-tuning subnetwork can be fine-tuned to minimize the adversarial loss and minimize the semantic control, thereby improving the semantic control accuracy of the image generation model in the process of generating the second image, as well as improving the image generation accuracy and image generation effect of the second image.
[0094] In an exemplary embodiment, the trained image generation model can be used to generate images. Figure 2 As shown, the process involved in the image generation method using the trained image generation model may include steps S201 to S203.
[0095] Step S201: Acquire a first image and image annotation text corresponding to the first image.
[0096] For example, the first image may be pre-set or may be set by the user. For example, a plurality of preset images may be pre-set in the computer device. The user may select one or more preset images from the plurality of preset images as the first image, so as to subsequently generate the corresponding second image based on the first image. For another example, the user may also input one or more preset images obtained by the user into the computer device, so that the computer device determines that the received preset image is the first image. Of course, this is not limited to this, and is not limited here.
[0097] When the first image is acquired, the computer device may perform image annotation processing on the first image to obtain image annotation text corresponding to the first image.
[0098] Taking the auto insurance business in the field of financial technology as an example, the first image may include a vehicle image or a non-vehicle image, which is not limited here. When the first image is a vehicle image, the image annotation text corresponding to the vehicle image includes, for example, "car". When the first image is a non-vehicle image, the image annotation text corresponding to the non-vehicle image includes, for example, "no car". Of course, it is not limited to this. For example, the computer device can also identify whether the vehicle image and the non-vehicle image contain pedestrians, so as to add text such as "pedestrians" and "no pedestrians" to the corresponding image annotation text. This is not limited here.
[0099] Taking the medical insurance business in the medical and health field as an example, the first image may include a disease image or a non-disease image, which is not limited here. In the case where the first image is a disease image, the image annotation text corresponding to the disease image includes, for example, the disease type corresponding to the disease image. In the case where the first image is a non-disease image, the image annotation text corresponding to the non-disease image includes, for example, no corresponding disease type. Of course, it is not limited to this. For example, the computer device can also identify whether the disease image and the non-disease image contain environmental areas, so as to add the corresponding environment type, no specific environment type and other texts in the corresponding image annotation text. This is not limited here.
[0100] Step S202: Obtain semantic prompt text.
[0101] For example, the semantic prompt text may be pre-set or may be set by the user. For example, a plurality of preset prompt texts may be pre-set in the computer device. The user may select one or more preset prompt texts from the plurality of preset prompt texts as semantic prompt texts, so as to subsequently generate a corresponding second image based on the semantic prompt text. For another example, the user may input one or more preset prompt texts into the computer device, so that the computer device determines that the received preset prompt text is a semantic prompt text. Of course, this is not limited to this, and no limitation is made here.
[0102] Taking the auto insurance business in the field of financial technology as an example, the semantic prompt text may include, for example, "adding the types of cars covered by the auto insurance product, the auto insurance scenarios applicable to the auto insurance product, adding the corresponding auto insurance product marketing statements to the image, and deleting the image area of the person in the image". Of course, the semantic prompt text is not limited to this and is not limited here.
[0103] Taking the medical insurance business in the medical and health field as an example, the semantic prompt text may include, for example, "adding the types of diseases covered by medical insurance, the judgment result of whether medical insurance can cover the types of diseases indicated in the image, the customer categories to which medical insurance applies, and adding corresponding medical product marketing statements to the image". Of course, the semantic prompt text is not limited to this and is not limited here.
[0104] Step S203: Based on the image generation model, the semantic prompt text and the image annotation text, perform image semantic control processing on the first image to obtain a second image corresponding to the first image.
[0105] For example, a computer device can input the acquired first image, the image annotation text corresponding to the first image, and the semantic prompt text into an image generation model, so that the image generation model performs image semantic control processing on the first image in response to the semantic prompt text and the image annotation text to obtain a corresponding second image.
[0106] Taking the auto insurance business in the field of financial technology, the semantic prompt text includes "adding the type of car covered by the auto insurance product, the auto insurance scenario applicable to the auto insurance product, adding the corresponding auto insurance product marketing statement to the image, and deleting the image area of the person in the image" as an example. In the case where the first image is a vehicle image, the image annotation text corresponding to the vehicle image includes, for example, a car and a pedestrian. The image generation model can determine the need to add the type of car covered by the auto insurance product, the auto insurance scenario applicable to the auto insurance product, and the auto insurance product marketing statement to the vehicle image based on the instructions of the semantic prompt text and the image annotation text, and delete the image area of the person where the pedestrian is located from the vehicle image. In the case where the first image is a non-vehicle image, the image annotation text corresponding to the non-vehicle image includes, for example, a non-vehicle and a pedestrian. The image generation model can determine the need to add the car, the type of car covered by the auto insurance product, the auto insurance scenario applicable to the auto insurance product, and the auto insurance product marketing statement to the non-vehicle image based on the instructions of the semantic prompt text and the image annotation text, and delete the image area of the person where the pedestrian is located from the non-vehicle image. Of course, it is not limited to this, and no limitation is made here.
[0107] Taking the medical insurance business in the field of medical health, the semantic prompt text includes "adding the types of diseases covered by medical insurance, the judgment results of whether medical insurance can cover the types of diseases indicated in the image, the customer categories to which medical insurance applies, and adding corresponding medical product marketing statements to the image" as an example. In the case where the first image is a disease image, the image annotation text corresponding to the disease image includes, for example, the disease type corresponding to the disease image. The image generation model can determine the types of diseases covered by medical insurance, the judgment results of whether medical insurance can cover the types of diseases indicated in the image, the customer categories to which medical insurance applies, and the medical product marketing statements to be added to the disease image based on the semantic prompt text and the instructions of the image annotation text. In the case where the first image is a non-disease image, the image annotation text corresponding to the non-disease image includes, for example, that there is no corresponding disease type. The image generation model can determine the types of diseases covered by medical insurance, the customer categories to which medical insurance applies, and the medical product marketing statements to be added to the non-disease image based on the semantic prompt text and the instructions of the image annotation text. Of course, it is not limited to this and is not limited here.
[0108] In this way, the trained image generation model can be applied to one or more semantic prompt texts to perform image semantic control processing on the first image. Then, during the generation process of the second image corresponding to the first image, the trained image generation model can have better semantic control accuracy on the second image.
[0109] The training method of the image generation model provided in the above embodiment obtains a training data set, the training data set includes multiple first images and image annotation texts corresponding to each first image; obtains semantic prompt text; based on the image generation model, the semantic prompt text and the image annotation text, performs image semantic control processing on the first image to obtain a second image corresponding to the first image; determines the adversarial loss between the first image and the second image according to the first image and the second image; determines the semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image; and adjusts the model parameters of the image generation model according to the adversarial loss and the semantic control loss.
[0110] For example, in the process of training the image generation model, the model parameters of the image generation model can be adjusted according to the adversarial loss and the semantic control loss. Since the adversarial loss is determined according to the first image and the second image, the image generation model can evaluate whether the second image is similar to the first image according to the adversarial loss. Correspondingly, since the semantic control loss is determined according to the semantic prompt text and the second image, the image generation model can evaluate whether the object contained in the second image is consistent with the object indicated by the semantic prompt text according to the semantic control loss. Based on this, when the model parameters of the image generation model are adjusted according to the adversarial loss and the semantic control loss, the image generation model can perform image semantic control processing on the first image according to the semantic prompt text to obtain a second image that is similar to the first image and contains objects that are consistent with the objects indicated by the semantic prompt text, which is conducive to improving the semantic control accuracy of the image generation model in the process of image generation, so as to improve the training effect of the image generation model.
[0111] The trained image generation model obtained by the training method of the image generation model can be applied to the fields of financial technology, medical health, etc. In some embodiments, for the auto insurance business in the field of financial technology, in response to the marketing needs of the auto insurance business for auto insurance products, for example, a large number of auto insurance product marketing images related to cars need to be generated, then the trained image generation model can be combined with the first image provided by the auto insurance business and the semantic prompt text related to the marketing of auto insurance products to generate a large number of second images similar to the first image and containing objects that match the objects indicated by the semantic prompt text. Based on the generation of the second image, the second image can be used for the marketing of auto insurance products, which is conducive to improving the convenience of marketing of auto insurance products. In other embodiments, for the medical insurance business in the field of medical health, in response to the marketing needs of the medical insurance business for medical insurance products, for example, a large number of medical insurance product marketing images related to medical care need to be generated, then the trained image generation model can be combined with the first image provided by the medical insurance business and the semantic prompt text related to the marketing of medical insurance products to generate a large number of second images similar to the first image and containing objects that match the objects indicated by the semantic prompt text. Based on the generation of the second image, the second image can be used for the marketing of medical insurance products, which is conducive to improving the convenience of marketing of medical insurance products. Of course, it is not limited to this and no limitation is made here.
[0112] See also Figure 3 , Figure 3 This is a schematic block diagram of a training device for an image generation model provided in an embodiment of the present application. The training device can be configured in a server or a computer device to execute the aforementioned training method for the image generation model.
[0113] like Figure 3 As shown, the training device of the image generation model includes: a first acquisition module 110, a second acquisition module 120, a processing module 130, a first loss determination module 140, a second loss determination module 150 and a model training module 160.
[0114] A first acquisition module 110 is used to acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images;
[0115] The second acquisition module 120 is used to acquire the semantic prompt text;
[0116] A processing module 130, configured to perform image semantic control processing on the first image based on an image generation model, the semantic prompt text, and the image annotation text, to obtain a second image corresponding to the first image;
[0117] A first loss determination module 140, configured to determine an adversarial loss between the first image and the second image according to the first image and the second image;
[0118] A second loss determination module 150, configured to determine a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image;
[0119] The model training module 160 is used to adjust the model parameters of the image generation model according to the adversarial loss and the semantic control loss.
[0120] It should be noted that technicians in the relevant field can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0121] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0122] Exemplarily, the above method and apparatus may be implemented in the form of a computer program, which may be run on a computer device.
[0123] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server or an electronic device.
[0124] like Figure 4 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a storage medium and an internal memory.
[0125] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any training method for an image generation model.
[0126] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0127] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any training method for the image generation model.
[0128] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0129] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0130] In one embodiment, the processor is used to execute a computer program and implement the following steps when executing the computer program:
[0131] Acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images;
[0132] Get semantic hint text;
[0133] Based on the image generation model, the semantic prompt text and the image annotation text, performing image semantic control processing on the first image to obtain a second image corresponding to the first image;
[0134] Determine, according to the first image and the second image, an adversarial loss between the first image and the second image;
[0135] Determining a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image;
[0136] According to the adversarial loss and the semantic control loss, model parameters of the image generation model are adjusted.
[0137] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the training of the image generation model described above can refer to the corresponding process in the aforementioned image generation model training method embodiment, and will not be repeated here.
[0138] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The method implemented when the computer program is executed by a processor can refer to the various embodiments of the training method of the image generation model of the present application.
[0139] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.
[0140] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0141] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0142] It should also be understood that the non-Company software tools or components appearing in the embodiments of the present application are merely examples and do not represent actual use.
[0143] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A training method for an image generation model, characterized in that: include: Acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images; Get semantic hint text; Based on the image generation model, the semantic prompt text and the image annotation text, performing image semantic control processing on the first image to obtain a second image corresponding to the first image; Determine, according to the first image and the second image, an adversarial loss between the first image and the second image; Determining a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image; According to the adversarial loss and the semantic control loss, model parameters of the image generation model are adjusted.
2. The training method according to claim 1, characterized in that: The determining, according to the first image and the second image, a confrontation loss between the first image and the second image includes: Based on the perceptual feature extraction network of the image generation model, performing perceptual feature extraction processing on the first image and the second image respectively to obtain a first perceptual feature corresponding to the first image and a second perceptual feature corresponding to the second image; determining a perceptual loss between the first perceptual characteristic and the second perceptual characteristic; Based on the residual network of the image generation model, semantic classification processing is performed on the first image and the second image respectively to obtain a first classification result corresponding to the first image and a second classification result corresponding to the second image; determining a classification loss between the first classification result and the second classification result; The adversarial loss is determined according to the sum of the perceptual loss and the classification loss.
3. The training method according to claim 1, characterized in that: The determining, according to the semantic hint text and the second image, a semantic control loss between the semantic hint text and the second image includes: Based on the text / image encoder of the image generation model, respectively perform embedding feature extraction processing on the semantic prompt text and the second image to obtain a first text embedding feature corresponding to the semantic prompt text and a second image embedding feature corresponding to the second image; The semantic control loss is determined according to a distance between the first text embedding feature and the second image embedding feature.
4. The training method according to any one of claims 1 to 3, characterized in that: The performing image semantic control processing on the first image based on the image generation model, the semantic prompt text and the image annotation text to obtain a second image corresponding to the first image includes: Based on the text / image encoder of the image generation model, respectively perform embedding feature extraction processing on the semantic prompt text, the image annotation text and the first image to obtain a first text embedding feature corresponding to the semantic prompt text, a second text embedding feature corresponding to the image annotation text and a first image embedding feature corresponding to the first image; Based on the stable diffusion network of the image generation model and the semantic fine-tuning subnetwork set in the stable diffusion network, the first image embedding feature is subjected to image semantic control processing according to the first text embedding feature and the second text embedding feature to obtain a second image corresponding to the first image.
5. The training method according to claim 4, characterized in that: The adjusting the model parameters of the image generation model according to the adversarial loss and the semantic control loss includes: According to the adversarial loss and the semantic control loss, the preset network parameters of the semantic fine-tuning subnetwork are adjusted, and the preset network parameters of the subnetworks other than the semantic fine-tuning subnetwork in the stable diffusion network are kept unchanged.
6. The training method according to any one of claims 1 to 3, characterized in that: The step of obtaining a training data set includes: Get multiple preset images; Eliminate preset images that do not meet the preset image screening condition from a plurality of preset images to obtain a first image; Perform image annotation processing on the first image to obtain image annotation text corresponding to the first image.
7. The training method according to claim 6, characterized in that: After eliminating preset images that do not meet the preset image screening condition from the plurality of preset images to obtain the first image, the method further includes: Performing image size conversion processing on the first image to obtain the first image in different image sizes; The performing image annotation generation processing on the first image to obtain image annotation text corresponding to the first image includes: Image annotation generation processing is performed on the first images at different image sizes respectively to obtain image annotation texts corresponding to the first images at different image sizes.
8. A training device for an image generation model, characterized in that: The training device comprises: A first acquisition module is used to acquire a training data set, where the training data set includes a plurality of first images and image annotation texts corresponding to each of the first images; The second acquisition module is used to acquire the semantic prompt text; A processing module, configured to perform image semantic control processing on the first image based on an image generation model, the semantic prompt text, and the image annotation text, to obtain a second image corresponding to the first image; A first loss determination module, configured to determine an adversarial loss between the first image and the second image according to the first image and the second image; A second loss determination module, configured to determine a semantic control loss between the semantic prompt text and the second image according to the semantic prompt text and the second image; A model training module is used to adjust the model parameters of the image generation model according to the adversarial loss and the semantic control loss.
9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the training method of the image generation model as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the training method of the image generation model as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Method and apparatus for training image generation model, and device and storage medium
WO2026152902A1