Training method for image generation model, and related apparatus and medium
By constructing an image generation model and training it using background template images, noisy reference images, and image description information, combined with denoising network control information, the problem of low object replacement accuracy in existing technologies is solved, and higher image generation accuracy is achieved.
Patent Information
- Application Number
- PCT/CN2025/100081
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2025-06-10
- Publication Date
- 2026-02-12
AI Technical Summary
Existing technologies, when replacing objects in a background image, are limited by factors such as angle, lighting, and filters, making it difficult to collect sufficient training data, resulting in low accuracy of the generated target image.
By constructing an image generation model, using background template images, noisy reference images, image description information, and sample object images, the model generates sample stitched image features. This feature is then combined with denoising network control information for training, thereby improving the accuracy of the model generation.
The accuracy of the image generation model in replacing objects has been enhanced, ensuring consistency between the target image and the background image, and improving the generation effect.
Smart Images

Figure CN2025100081_12022026_PF_FP_ABST
Abstract
Description
Method for training image generation model, related device and medium
[0001] The present application claims priority to the Chinese patent application No. 2024110606399, filed on August 5, 2024, and titled "Method for training image generation model, related device and medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to the field of artificial intelligence, and in particular, to a method for training an image generation model, related device and medium. BACKGROUND
[0003] Currently, in various business scenarios such as video production and virtual display, it is often necessary to replace background images and object images to create personalized images. For example, when producing an object image of item A, item B in a background image is replaced by item A. For this purpose, in the related art, a neural network model is often used to remove a first object in a specified background image, and a second object with the same orientation and shooting angle as the first object in the background image is produced, and the second object is embedded into the background image in the form of a map, so that the second object is displayed in the specified background image.
[0004] However, the effect of the target image generated by the above method is often limited by the angle, lighting filter and other factors of the target object (i.e. the second object) and the background image. In actual model training, it is often difficult to collect enough training data that meets the expectations, which affects the model training effect, resulting in the situation that the target image generated by the model does not meet the requirements (e.g. the lighting and filter of the target object in the image cannot be consistent with the background image), thereby leading to low accuracy of the target image generated by the model. SUMMARY
[0005] The present disclosure provides a method for training an image generation model, related device and medium, which can improve the accuracy of the target image generated by the image generation model.
[0006] According to an aspect of the present disclosure, a method for training an image generation model is provided, which is executed by an electronic device, and the training method comprises:
[0007] obtaining a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image comprising a sample object, and the noise reference image is obtained by adding noise to a reference object in the background template image and replacing the reference object with the sample object;
[0008] determine sample splicing image features based on the noise reference image, the background template image, and a template mask image, wherein the template mask image is obtained by masking the reference object in the background template image;
[0009] determine denoising network control information of the image generation model based on contour features of the reference object in the background template image, the image description information, and the sample object image;
[0010] perform noise prediction based on the sample splicing image features and the denoising network control information through the image generation model to obtain a noise prediction result corresponding to the noise reference image;
[0011] train the image generation model based on comparison results of the noise reference image and the corresponding noise prediction result in each of the plurality of image-text sample pairs.
[0012] According to an aspect of the present disclosure, an image generation method is provided, which is executed by an electronic device, and the image generation method comprises:
[0013] obtain an object image including a first object, a background image, and description information, wherein the background image contains a second object, and the description information is used to describe replacement from the second object to the first object;
[0014] determine splicing image features based on the background image, a preset noise image, and a mask image, wherein the mask image is obtained by masking the second object in the background image;
[0015] determine denoising control information of an image generation model based on contour features of the second object in the background image, the description information, and the object image, wherein the image generation model is generated according to the training method of the image generation model described above;
[0016] perform image generation through the image generation model based on the splicing image features and the denoising control information to obtain a target image, wherein the target image is used to indicate a result of replacing the first object of the object image with the second object in the background image.
[0017] According to an aspect of the present disclosure, an image generation model training device is provided, which comprises:
[0018] The first obtaining unit is configured to obtain a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image comprising a sample object, and the noise reference image is obtained by adding noise to a reference object obtained by replacing the sample object in the background template image;
[0019] The first determining unit is configured to determine a sample splicing image feature based on the noise reference image, the background template image, and a template mask image, wherein the template mask image is obtained by masking the reference object in the background template image;
[0020] The second determining unit is configured to determine denoising network control information of the image generation model based on contour features of the reference object in the background template image, the image description information, and the sample object image;
[0021] The predicting unit is configured to perform noise prediction based on the sample splicing image feature and the denoising network control information by using the image generation model to obtain a noise prediction result corresponding to the noise reference image;
[0022] The training unit is configured to train the image generation model based on a comparison result of the noise reference image in the plurality of image-text sample pairs and the noise prediction result corresponding thereto.
[0023] According to an aspect of the present disclosure, an image generation device is provided, which comprises:
[0024] The second obtaining unit is configured to obtain an object image comprising a first object, a background image comprising a second object, and description information for describing replacement from the second object to the first object;
[0025] The third determining unit is configured to determine a splicing image feature based on the background image, a preset noise image, and a mask image, wherein the mask image is obtained by masking the second object in the background image;
[0026] The fourth determining unit is configured to determine denoising control information of an image generation model based on contour features of the second object in the background image, the description information, and the object image, wherein the image generation model is generated by the above-mentioned image generation model training method;
[0027] An image generation unit is configured to generate an image based on the stitching image feature and the denoising control information by using the image generation model, and obtain a target image, where the target image is used to indicate a result of replacing the second object in the background image with the first object in the object image.
[0028] According to an aspect of the present disclosure, an electronic device is provided, which includes a memory storing a computer program and a processor, wherein the processor implements the training method of the image generation model or the image generation method when executing the computer program.
[0029] According to an aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program, wherein the computer program is executed by a processor to implement the training method of the image generation model or the image generation method.
[0030] According to an aspect of the present disclosure, a computer program product is provided, which includes a computer program, wherein the computer program is read and executed by a processor of an electronic device, so that the electronic device executes the training method of the image generation model or the image generation method.
[0031] In the embodiments of the present disclosure, when training the image generation model, a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including a sample object are used to construct a sample image-text pair. The noise reference image is obtained by adding noise to a reference object in the background template image. This way can use the reference noise reference image to make it more realistic. Then, the image features of the background template image, the object mask image corresponding to the background template image, and the noise reference image are integrated into sample splicing image features. The sample splicing image features are used as input for model training. Since the sample splicing image features have multiple image feature information, they can be used as training data to enrich the information amount of the training data. In addition, the present disclosure also introduces image description information (which can indicate how the background and object in the noise reference image are), contour features of the reference object in the background template image (which can reflect the action and posture of the reference object), and a sample object image including a sample object (which can reflect the contour and appearance of the sample object) to generate denoising network control information together, so that the denoising network control information has multiple constraint conditions. Further, the denoising network control information and the sample splicing image features are input into the image generation model, so that the image generation model is limited by the denoising network control information when predicting noise from the sample splicing image features, and the noise prediction process is corrected according to the denoising network control information, so that the noise prediction result of the noise reference image output by the image generation model is more accurate. Finally, the image generation model is trained by comparing the difference between the noise prediction result and the noise contained in the noise reference image, so that the image generation model that meets the training requirements is obtained. This way can improve the accuracy of the model generated image.
[0032] Other features and advantages of the present disclosure will be set forth in the following description, and in part will become apparent from the description, or can be learned by practice of the present disclosure. The objects and other advantages of the present disclosure can be achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0033] The accompanying drawings are included to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. The drawings, together with the embodiments of the present disclosure, are used to explain the technical solutions of the present disclosure, and do not constitute a limitation on the technical solutions of the present disclosure.
[0034] FIG. 1 is a system architecture diagram of a training method and an image generation method of an image generation model according to an embodiment of the present disclosure;
[0035] FIG. 2A is one of the schematic diagrams of the image generation method applied in a video production scenario according to an embodiment of the present disclosure;
[0036] FIG. 2B is another of the schematic diagrams of the image generation method applied in a video production scenario according to an embodiment of the present disclosure;
[0037] FIG. 2C is a third of the schematic diagrams of the image generation method applied in a video production scenario according to an embodiment of the present disclosure;
[0038] FIG. 3 is a flowchart of a training method of an image generation model according to an embodiment of the present disclosure;
[0039] FIG. 4 is a flowchart of determining a sample object image according to an embodiment of the present disclosure;
[0040] FIG. 5 is a flowchart of determining a noise reference image according to an embodiment of the present disclosure;
[0041] FIG. 6 is a schematic diagram of an implementation process of determining a noise reference image according to an embodiment of the present disclosure;
[0042] FIG. 7 is a flowchart of generating a sample splicing image feature according to an embodiment of the present disclosure;
[0043] FIG. 8 is a flowchart of determining a template mask image according to an embodiment of the present disclosure;
[0044] FIG. 9A is a schematic diagram of an implementation process of determining a template mask image according to an embodiment of the present disclosure;
[0045] FIG. 9B is another schematic diagram of an implementation process of determining a template mask image according to an embodiment of the present disclosure;
[0046] FIG. 10 is a flowchart of generating denoising network control information according to an embodiment of the present disclosure;
[0047] FIG. 11 is a flowchart of determining a contour feature of a reference object in a background template image according to an embodiment of the present disclosure;
[0048] FIG. 12 is a schematic diagram of an implementation process of key point extraction when determining a contour feature according to an embodiment of the present disclosure;
[0049] FIG. 13 is a flowchart of generating an image description embedding vector according to an embodiment of the present disclosure;
[0050] FIG. 14 is a schematic diagram of an implementation process of generating an image description embedding vector according to an embodiment of the present disclosure;
[0051] FIG. 15 is a flowchart of generating denoising network control information according to an embodiment of the present disclosure;
[0052] Figure 16 is a schematic diagram of the implementation process of generating denoised network control information according to an embodiment of the present disclosure;
[0053] Figure 17 is a flowchart of generating noise prediction results according to an embodiment of the present disclosure;
[0054] Figure 18 is a schematic diagram of the process of generating noise prediction results according to an embodiment of the present disclosure;
[0055] Figure 19 is a flowchart of a noise reduction process according to an embodiment of the present disclosure;
[0056] Figure 20 is a schematic diagram illustrating the implementation process of the downsampling process during denoising processing according to an embodiment of the present disclosure;
[0057] Figure 21 is a schematic diagram illustrating the implementation process of the upsampling process during denoising according to an embodiment of the present disclosure;
[0058] Figure 22 is a flowchart of a training image generation model according to an embodiment of the present disclosure;
[0059] Figure 23 is a flowchart of determining a sub-loss function according to an embodiment of the present disclosure;
[0060] Figure 24 is a schematic diagram illustrating the implementation details of a training image generation model according to an embodiment of the present disclosure;
[0061] Figure 25 is a flowchart of an image generation method according to an embodiment of the present disclosure;
[0062] Figure 26 is a flowchart of determining target stitched image features according to an embodiment of the present disclosure;
[0063] Figure 27 is a flowchart of generating a target image according to an embodiment of the present disclosure;
[0064] Figure 28 is a simplified schematic diagram of an embodiment of an image generation method according to the present disclosure;
[0065] Figure 29 is a schematic diagram illustrating the implementation details of an image generation method according to an embodiment of the present disclosure;
[0066] Figure 30 is a block diagram of a training apparatus for an image generation model according to an embodiment of the present disclosure;
[0067] Figure 31 is a block diagram of an image generation apparatus according to an embodiment of the present disclosure;
[0068] Figure 32 is a terminal structure diagram of a training method for an image generation model according to an embodiment of the present disclosure;
[0069] Figure 33 is a server structure diagram of a training method for an image generation model according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0070] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and not intended to limit the present disclosure.
[0071] Before the embodiments of the present disclosure are further described in detail, the terms and phrases involved in the embodiments of the present disclosure are explained, and the terms and phrases involved in the embodiments of the present disclosure are applicable to the following explanations:
[0072] Cross Attention Control Module: a module used in deep learning to establish cross-attention mechanisms between multiple inputs. The Cross Attention Control Module can help the model automatically learn the correlation between inputs when processing multiple inputs, thereby improving the performance of the model.
[0073] The system architecture and scenarios to which the embodiments of the present disclosure are applied are described below.
[0074] FIG. 1 is a system architecture diagram to which a training method of an image generation model and an image generation method according to an embodiment of the present disclosure are applied. It includes an object terminal 140, the Internet 130, a gateway 120, an image processing server 110, an image database 150, and the like.
[0075] The object terminal 140 includes desktop computers, laptop computers, PDAs (Personal Digital Assistants), tablet computers, mobile phones, car-mounted terminals, home theater terminals, smart TVs, dedicated terminals, and the like in various forms. In addition, it can be a single device or a collection of multiple devices. The object terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data. The object terminal 140 includes an image processing system for receiving an object-selected background image and an object image, and submitting the object-selected background image and the object image to the image processing server 110 so that the image processing server 110 generates a target image with a target object in the object image as a foreground and the background image as a background.
[0076] The image processing server 110 refers to a computer system capable of providing certain services to the object terminal 140. Compared with the ordinary object terminal 140, the image processing server 110 has higher requirements in stability, security, performance, and the like. The image processing server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (for example, a virtual machine) of a high-performance computer, a combination of parts (for example, virtual machines) of multiple high-performance computers, or a cloud server, and the like. The image processing server 110 includes multiple types of services, and the implementation of each service of the image processing server 110 is often associated with some intermediate databases or storage media, and the like. The image processing server 110 is configured to generate a target image with a target object in an object image as a foreground and a background image as a background by using a trained image generation model, and the image database 150 is configured to store various images such as the target image, the object image, and the background image.
[0077] The gateway 120 is also referred to as an inter-network connector or a protocol converter. The gateway implements network interconnection at the transport layer and is a computer system or device acting as a conversion function. The gateway is a translator between two systems using different communication protocols, data formats or languages, or even having completely different architectures. Meanwhile, the gateway can also provide filtering and security functions. The messages sent by the object terminal 140 to the image processing server 110 need to be sent to the corresponding image processing server 110 through the gateway 120. The messages sent by the image processing server 110 to the object terminal 140 also need to be sent to the corresponding object terminal 140 through the gateway 120.
[0078] The embodiments of the present disclosure can be applied in various scenarios, such as the video production scenarios shown in FIGS. 2A-2C, and the like.
[0079] As shown in FIG. 2A, when the object needs to produce an image with the object A as a foreground and the scene of the background image 5 as a background in the video production process, the object logs in to the image processing system on the object terminal and enters the image processing process. At this time, a prompt field “please select a background image, upload a target object image, and input an image description” is displayed on the page, and an editing area for selecting a background image, an editing area for selecting a target object image, and an editing area for inputting an image description are provided. Based on this, the object selects “background image 5” in the editing area for selecting a background image; the object selects “D:\object image of object A\image 1+image 2” in the editing area for selecting a target object image; the object inputs “replace object B in the background image 5 with object A” in the editing area for inputting an image description, and clicks the “OK” button to determine that the background image 5, the image 1 and the image 2 of the object A are used to replace object B in the background image 5 with object A.
[0080] As shown in FIG. 2B, after the object clicks the "OK" button, a prompt window is displayed on the page, wherein the prompt window has a prompt field "Mask image corresponding to background image 5 and object contour image of object B are being generated, and the image generation model is called to perform image synthesis according to the mask image and the object contour image, please wait patiently... " to indicate to the object the image synthesis process of the image generation model.
[0081] As shown in FIG. 2C, when the image generation model is executed, the page displays a prompt field "Object B in background image 5 is replaced by object A, see the following figure. In addition, the generated image has been saved to the path: D:\Synthesized image folder\Video production.", and a comparison of background image 5 and the generated target image is displayed. Among them, background image 5 has bottle A (object B), and target image has bottle B (object A), and through the image generation model in the image processing system, bottle A in background image 5 is replaced by bottle B while keeping the background unchanged.
[0082] It should be noted that the image generation method of the embodiment of the present disclosure can be applied not only to the video production scene described above, but also to various application scenarios such as the product display scene in e-commerce and the personalized image production scene in social media.
[0083] The training method of the image generation model of the embodiment of the present disclosure will be described below.
[0084] According to one embodiment of the present disclosure, a training method of an image generation model is provided.
[0085] The training method of the image generation model is generally applied to business scenarios where a reference object (person, animal, object, etc.) in a fixed background image is replaced by a target object (person, animal, object, etc.), such as the video production scene, product display scene, etc. shown in FIGS. 2A-2C. The embodiment of the present disclosure provides a scheme for model training based on image description, object difference between the image to be generated and the background image, which can improve the accuracy of the image generation model in generating target images.
[0086] As shown in FIG. 3, the training method of the image generation model according to one embodiment of the present disclosure can be executed by an electronic device, which can be an image processing server or an object terminal shown in FIG. 1, and the training method of the image generation model can include:
[0087] Step 310, obtaining a plurality of image-text sample pairs;
[0088] Step 320, determining sample splicing image features based on the noise reference image, the background template image, and the template mask image.
[0089] Step 330, determining the denoising network control information of the image generation model based on the contour feature of the reference object in the background template image, the image description information, and the sample object image;
[0090] Step 340, performing noise prediction on the sample splicing image feature and the denoising network control information through the image generation model to obtain a noise prediction result corresponding to the noise reference image;
[0091] Step 350, training the image generation model based on the comparison result of the noise reference image and the noise prediction result corresponding thereto in each of the plurality of pairs of image-text samples.
[0092] The steps 310-350 are described in detail as follows.
[0093] In step 310, a plurality of pairs of image-text samples are obtained.
[0094] In the embodiments of the present disclosure, one pair of image-text samples serves as one training data, wherein each pair of image-text samples includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including a sample object.
[0095] The background template image refers to an image providing background content for the finally required generated image, and in the embodiments of the present disclosure, the sample object needs to be added to the background of the background template image, wherein the background template image also contains a reference object to be replaced by the sample object in the sample object image.
[0096] The sample object image including a sample object refers to an image reflecting the object characteristics (object posture, object action, object appearance, etc.) of the sample object, and each pair of image-text samples can include one or more sample object images.
[0097] The noise reference image is obtained by adding noise to the reference object in the background template image, that is, the sample object can be used to replace the reference object in the background template image to obtain a reference image, and then noise is iteratively added to the reference image to obtain the noise reference image.
[0098] The image description information corresponding to the noise reference image is used to indicate the conditional constraints on the denoising process. For example, the image description information can be to generate a person A in the background template image.
[0099] In step 320, the sample splicing image feature is determined based on the noise reference image, the background template image, and the template mask image.
[0100] The template mask image is obtained by masking a reference object in the background template image.
[0101] The sample spliced image feature is used to indicate a splicing result of respective features of the noise reference image, the background template image, and the template mask image.
[0102] In the implementation of this embodiment, first, the noise reference image, the background template image, and the template mask image are respectively mapped from a pixel space to a latent vector space, to obtain a vector feature corresponding to the noise reference image, a vector feature corresponding to the background template image, and a vector feature corresponding to the template mask image. Then, the vector feature corresponding to the noise reference image, the vector feature corresponding to the background template image, and the vector feature corresponding to the template mask image are spliced to obtain the sample spliced image feature.
[0103] In step 330, based on the contour feature of the reference object in the background template image, the image description information, and the sample object image, the denoising network control information of the image generation model is determined.
[0104] The contour feature of the reference object in the background template image is used to indicate a contour characteristic of an action, a posture, or the like of the reference object in the background template image.
[0105] The image generation model of the embodiment of the present disclosure is a neural network model constructed based on a diffusion model (stable diffusion, SD). The image generation model of the embodiment of the present disclosure often takes a background image, an object image, and information describing replacement of an object in the object image with an object in the background image as input, and outputs a composite image taking a scene of the background image as a background and the object in the object image as a foreground. The image generation model can realize embedding of an object into a specific position of a certain image, and meet the image generation demand in multiple scenarios.
[0106] The denoising network control information is used as a conditional constraint to assist the image generation model in image denoising.
[0107] In the implementation of this embodiment, the image description information can be encoded to obtain a corresponding image description feature vector, the sample object image can be subjected to feature extraction processing to obtain a corresponding sample object feature vector, and then the denoising network control information can be determined according to the contour feature of the reference object in the background template image, and the image description feature vector and the sample object feature vector.
[0108] For the sake of brevity, the specific process of determining the denoising network control information based on the contour feature of the reference object in the background template image, the image description information, and the sample object image of the embodiment of the present disclosure will be described in more detail hereinafter, and will not be described here.
[0109] In step 340, based on the sample spliced image feature and the denoising network control information, noise prediction is performed through the image generation model to obtain a noise prediction result corresponding to the noise reference image.
[0110] The noise prediction result is used to indicate the noise contained in the image feature generated through the image generation model and consistent with the noise reference image.
[0111] In the implementation of this embodiment, the sample spliced image feature can be taken as the object to be denoised of the image generation model, and the denoising network control information can be taken as the control information referenced when the image generation model performs denoising processing. When the image generation model performs denoising processing, the sample spliced image feature can be iteratively denoised based on the denoising network control information, so as to obtain the noise prediction result of the noise reference image in each round of denoising processing.
[0112] For the sake of brevity, the specific process of noise prediction based on the sample spliced image feature and the denoising network control information through the image generation model will be described in more detail below, and will not be described here.
[0113] In step 350, based on the comparison result of the noise reference image and the corresponding noise prediction result in each of the plurality of image-text sample pairs, the image generation model is trained.
[0114] In the implementation of this embodiment, according to the noise difference between the noise contained in the noise reference image of the image-text sample pair and the noise prediction result, the model parameters of the image generation model are adjusted, and the above steps 310-350 are repeated to continuously reduce the noise difference between the noise contained in the noise reference image of the image-text sample pair and the noise prediction result, until the noise difference between the noise contained in the noise reference image of the image-text sample pair and the noise prediction result meets the model training requirement, the updating of the model parameters of the image generation model is stopped, the model parameters at this time are taken as the final model parameters, and the image generation model with the final model parameters is taken as the trained image generation model.
[0115] By the steps 310-350, in training the image generation model, the background template image, the noise reference image, the image description information corresponding to the noise reference image, and the sample object image of the sample object are used to construct the image-text sample pair, wherein the noise reference image is obtained by adding noise to the reference image obtained by replacing the reference object in the background template image with the sample object. This way can use the reference image to make the reference image more realistic. Then, the image features of the background template image, the template mask image of the background template image (obtained by masking the reference image in the background template image), and the noise reference image are integrated into the sample splicing image feature, and the sample splicing image feature is used as the input of the model training. Since the sample splicing image feature has multiple image feature information, it can be used as training data to enrich the information amount of the training data. In addition, the image description information (which can indicate what the background and object in the noise reference image are), the contour feature of the reference object in the background template image (which can reflect the action, posture, etc. of the reference object), and the sample object image of the sample object (which can reflect the contour and appearance of the sample object) are introduced to generate the denoising network control information together, so that the denoising network control information has multiple constraint conditions. Further, the denoising network control information and the sample splicing image feature are input into the image generation model, so that the image generation model is limited by the denoising network control information when predicting the noise of the sample splicing image feature, and the noise prediction process is corrected according to the denoising network control information, so that the noise prediction result of the noise reference image output by the image generation model is more accurate. Finally, the image generation model is trained by comparing the noise prediction result and the noise difference of the noise contained in the noise reference image, so as to obtain the image generation model that meets the training requirements. This way can improve the accuracy of the target image generated by the model.
[0116] The above is a general description of steps 310-350. The specific implementation of steps 310, 320, 330, 340, and 350 will be described in detail below.
[0117] Step 310 will be described in detail below.
[0118] In step 310, a plurality of image-text sample pairs are obtained, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including a sample object, wherein the noise reference image is obtained by adding noise to the reference image obtained by replacing the reference object in the background template image with the sample object.
[0119] Referring to FIG. 4, in one embodiment, the sample object image of the sample object is determined by the following manner:
[0120] Step 410, obtaining a sample image including a sample object;
[0121] Step 420, performing image segmentation on the sample image based on a preset object segmentation model to obtain a sample segmentation image having the sample object;
[0122] Step 430, performing image enhancement on the sample segmentation image to obtain the sample object image.
[0123] The steps 410-430 are described in detail as follows.
[0124] In step 410, the sample image is an image including the sample object.
[0125] In the implementation of this embodiment, under the condition of authorized permission, a plurality of sample images including the sample object can be extracted from an existing image database, and video data having the sample object can also be extracted from an existing video database, and the video data is subjected to video frame segmentation, and each video frame having the sample object is taken as a sample image.
[0126] In step 420, the sample segmentation image refers to a local image having the sample object segmented from the sample image, and the sample segmentation image is a part of the sample image.
[0127] The object segmentation model can be a lightweight semantic segmentation model BiSeNet-v2 or PP_LiteSeg.
[0128] The object segmentation model is taken as an example of the semantic segmentation model PP_LiteSeg, and the object segmentation model includes an encoding module, a pyramid pooling module, a decoding module, and an attention fusion module. Specifically, first, the sample image is input to the encoding module of the object segmentation model, and the sample image is encoded by the encoding module to obtain a first sample image feature with an image size of one quarter of the sample image, a second sample image feature with an image size of one eighth of the sample image, a third sample image feature with an image size of one sixteenth of the sample image, and a fourth sample image feature with an image size of one thirty-second of the sample image. Then, the fourth sample image feature is subjected to feature pooling by the pyramid pooling module to obtain a sample image pooling feature. Further, the sample image pooling feature and the third sample image feature are subjected to feature fusion by the attention fusion module to obtain a first image fusion feature. Then, the first image fusion feature and the second sample image feature are subjected to feature fusion by the attention fusion module to obtain a second image fusion feature. Further, the second image fusion feature is subjected to decoding processing by the decoding module to obtain a sample segmentation image with a sample object.
[0129] In step 430, the image enhancement of the sample segmentation image includes, but is not limited to, contrast enhancement, brightness enhancement, sharpening, and noise reduction of the sample segmentation image. Specifically, first, when the sample segmentation image is subjected to image enhancement, the pixel value distribution of the sample segmentation image is expanded using a linear stretching or a logarithmic transformation method to make the light and dark areas of the sample image segmentation image more distinct, so as to achieve contrast enhancement of the sample segmentation image. Then, after the contrast enhancement, the brightness values of all the pixel points of the sample segmentation image are adjusted to improve the overall brightness of the sample segmentation image. Further, after the brightness adjustment, the sample segmentation image is smoothed and the noise in the sample segmentation image is reduced using a Gaussian filter or a median filter. Finally, after the noise reduction, the sample segmentation image is subjected to edge enhancement by a Laplace operator to achieve sharpening processing of the sample segmentation image, and a sample object image is obtained.
[0130] The advantage of this embodiment is that the sample image including the sample object is subjected to image segmentation, and the local image with the sample object is segmented from the sample image, which can better eliminate the interference of irrelevant image information and improve the image quality of the sample object image used for training. Further, various image enhancement processing such as contrast enhancement, brightness enhancement, sharpening processing, and noise reduction processing is performed on the sample segmentation image, which can further improve the image quality of the sample object image.
[0131] Please refer to FIG. 5, in one embodiment, the noise reference image is determined by the following way:
[0132] Step 510, determining a reference image in which a reference object in the background template image is replaced by the sample object;
[0133] Step 520, generating a random number subject to a Gaussian distribution based on a predetermined random number generation model;
[0134] Step 530, adding the random number to the pixel value of each pixel point in the reference image to obtain a noise reference image.
[0135] The steps 510-530 are described in detail below.
[0136] In step 510, the reference image is the image data obtained after the reference object in the background template image is replaced by the sample object.
[0137] In the implementation of this embodiment, the image editing software can be used to replace the reference object in the background template image with the sample object, so that a replaced image is obtained as the reference image in which the reference object in the background template image is replaced by the sample object. The image editing software can be Adobe Photoshop or the like.
[0138] In step 520, the predetermined random number generation model refers to a random number generator, and the random number is a randomly generated number. In this embodiment, the random number subject to the Gaussian distribution can be repeated.
[0139] In the implementation of this embodiment, the predetermined random number generation model has a library function (numpy.random.normal() function). Specifically, according to the mean of 0 and the standard deviation of 1 to be followed by the Gaussian distribution, a random number subject to the Gaussian distribution is generated through the library function of the predetermined random number generation model. The random number subject to the Gaussian distribution can be expressed in the form of a random noise image, and the image size of the random noise image is the same as that of the reference image.
[0140] In step 530, for each pixel point in the reference image, first, the pixel value of the pixel point in the reference image is determined, and the random number corresponding to the pixel point in the random noise image is determined according to the position of the pixel point in the reference image, and the random number is taken as the noise value corresponding to the pixel point. Then, the pixel value of the pixel point and the corresponding noise value are added to obtain the noisy pixel value of the pixel point. Finally, according to the noisy pixel value of each pixel point, the noise reference image is obtained.
[0141] As shown in FIG. 6, it is a specific implementation process diagram of superimposing random noise on the reference image. The reference image is an 8x8 image with 64 pixel points. The random noise image corresponding to the random number obeying Gaussian distribution is also an 8x8 image with 64 pixel points. Specifically, for each pixel point of the reference image, the pixel value of each pixel point is determined, wherein the pixel value of each pixel point includes 1, 2, 3, 4, 5 or 6. Then, for each pixel point, the noise value corresponding to each pixel point in the random noise image is determined, wherein the noise value includes 0, 1, 2, 3, 4, 5 or 6. Further, for each pixel point, the pixel value and the corresponding noise value are added to obtain the noise pixel value of each pixel point. For example, for the first row of pixel points in the reference image, the pixel value of the first pixel point is 1, and the noise value is 5, so the noise pixel value is 1+5=6. The pixel value of the second pixel point is 1, and the noise value is 1, so the noise pixel value is 1+1=2. The pixel value of the third pixel point is 1, and the noise value is 4, so the noise pixel value is 1+4=5; and so on, until the noise pixel value of the last pixel point in the last row is determined, to generate the noise reference image.
[0142] The advantage of this embodiment is that the random noise conforming to Gaussian distribution is added to the reference image obtained by replacing the reference object in the background template image with the sample object, the environmental noise is introduced in the ideal reference image, the noise reference image obtained finally is more consistent with the real situation, thereby improving the authenticity and accuracy of the noise reference image, and the noise reference image has better referenceability.
[0143] The step 320 is described in detail below.
[0144] In the step 320, based on the noise reference image, the background template image, and the template mask image, the sample splicing image feature is determined, wherein the template mask image is obtained by masking the reference object in the background template image.
[0145] Referring to FIG. 7, in one embodiment, the step 320 specifically includes but is not limited to the following steps 710-740:
[0146] Step 710, performing first encoding processing on the noise reference image to obtain reference image encoding features;
[0147] Step 720, performing second encoding processing on the background template image to obtain background image encoding features;
[0148] Step 730, performing third encoding processing on the template mask image to obtain mask image encoding features;
[0149] Step 740, splicing the reference image encoding feature, the background image encoding feature, and the mask image encoding feature to obtain a sample spliced image feature.
[0150] The steps 710-740 are described in detail below.
[0151] In step 710, the reference image encoding feature is used to indicate the result of converting the noise reference image from the data space to the pixel space.
[0152] In the implementation of this embodiment, a preset image encoder can be used to perform first encoding processing on the noise reference image, convert the noise reference image from the pixel space to the latent vector space, capture image key information of the noise reference image, and obtain the reference image encoding feature, where the image key information of the noise reference image includes but is not limited to texture information, edge information, and corner information of the noise reference image.
[0153] In step 720, the background image encoding feature is used to indicate the result of converting the background template image from the data space to the pixel space.
[0154] In the implementation of this embodiment, the specific process of step 720 is similar to the specific process of step 710 described above. The difference lies in that the images to be encoded are different, and the encoder parameters of the image encoders used are different. To save space, no longer tedious.
[0155] In step 730, the mask image encoding feature is used to indicate the result of converting the template mask image from the data space to the pixel space.
[0156] In the implementation of this embodiment, the specific process of step 730 is similar to the specific process of step 710 described above. The difference lies in that the images to be encoded are different, and the encoder parameters of the image encoders used are different. To save space, no longer tedious.
[0157] In step 740, the reference image encoding feature, the background image encoding feature, and the mask image encoding feature are spliced to splice the reference image encoding feature, the background image encoding feature, and the mask image encoding feature into a vector feature with a larger number of feature channels, and the spliced vector feature is determined as a sample spliced image feature.
[0158] For example, the feature dimensions of the reference image encoding feature, the background image encoding feature, and the mask image encoding feature are all W*H*4, where W is the feature length of the reference image encoding feature, the background image encoding feature, and the mask image encoding feature, H is the feature height of the reference image encoding feature, the background image encoding feature, and the mask image encoding feature, and 4 is the feature channel number of the reference image encoding feature, the background image encoding feature, and the mask image encoding feature. After the above feature splicing operation, the feature dimension of the sample spliced image feature obtained is W*H*12.
[0159] The advantage of this embodiment is that the image information of the noise reference image, the background template image, and the template mask image is all converted from the pixel space to the latent vector space, and the image information (reference image encoding feature, background image encoding feature, and mask image encoding feature) of the noise reference image, the background template image, and the template mask image in the latent vector space is spliced to form a sample spliced image feature having noise-containing reference image information, template image information, and reference object mask information. Further, the sample spliced image feature is used as the input of the image generation model during model training, so that the image generation effect information, the background template information, and the reference object information in the background template are fused in the input data of the model, which can better improve the richness and comprehensiveness of the feature information of the sample spliced image feature, and is conducive to the learning and mining of the model on various image information, thereby improving the image generation accuracy of the model.
[0160] Please refer to FIG. 8. In one embodiment, the template mask image is determined by the following method:
[0161] Step 810: determining an object contour region of the reference object in the background template image;
[0162] Step 820: replacing the pixel values of each pixel point in the object contour region in the background template image with a first value, and replacing the pixel values of each pixel point outside the object contour region with a second value to obtain the template mask image.
[0163] The steps 810-820 are described in detail below.
[0164] In step 810, the object contour region is used to indicate the minimum circumscribed rectangular region corresponding to the object contour of the reference object.
[0165] In the embodiment, first, the pixels representing the reference object are determined in the background template image. Then, a two-dimensional coordinate system is constructed with the top-left pixel in the background template image as the origin, the height direction of the background template image as the vertical axis, the width direction of the background template image as the horizontal axis, and the distance between two adjacent pixels as a unit length. Further, the coordinate data of the pixels representing the reference object are determined based on the constructed two-dimensional coordinate system. The maximum horizontal coordinate x max , the minimum horizontal coordinate x min , the maximum vertical coordinate y max , and the minimum vertical coordinate y min are selected from the coordinate data of the pixels. Finally, the minimum circumscribed rectangular region corresponding to the object contour of the reference object is determined based on the maximum horizontal coordinate, the minimum horizontal coordinate, the maximum vertical coordinate, and the minimum vertical coordinate, and the minimum circumscribed rectangular region is determined as the object contour region of the reference object. The four end point coordinates of the minimum circumscribed rectangular region are (x max , y max ), (x max , y min ), (x min , y max ), and (x min , y min ).
[0166] In step 820, the first value refers to the value 1 that makes the pixel white, and the second value refers to the value 0 that makes the pixel black.
[0167] In the embodiment, first, the background template image is divided into the object contour region and other regions outside the object contour region. Then, for each pixel in the object contour region, the pixel value of the pixel is replaced by the first value, so that the object contour region appears white. For each pixel outside the object contour region, the pixel value of the pixel is replaced by the second value, so that the object contour region appears black, thereby obtaining the template mask image according to the change of the pixel value of each pixel in the background template image. The template mask image is often represented as a black and white image.
[0168] As shown in FIG. 9A, it is a brief schematic of the mask processing at the pixel level. Specifically, in the background template image with the image size of 8x8, the minimum circumscribed rectangle region corresponding to the object contour of the reference object is a 4x4 image region, wherein the object contour region contains 3 pixel points with the pixel value of 9, 2 pixel points with the pixel value of 8, 2 pixel points with the pixel value of 1, 2 pixel points with the pixel value of 4, and 7 pixel points with the pixel value of 2. Based on this, all the pixel values of the 16 pixel points are replaced with 1, and the pixel values of the pixel points other than the 16 pixel points in the background template image are replaced with 0, to obtain a template mask image in which the pixel values of the pixel points are composed of 0 or 1. In the template mask image, the object contour region of the reference object is white, and the other parts are black.
[0169] As shown in FIG. 9B, it is a brief schematic of the mask processing at the image level. Specifically, the background template image contains a reference object (a person) and several other objects (a pentagram and a polygon). Based on this, the minimum circumscribed rectangle corresponding to the contour of the reference object (the person) is first determined, the pixel values of the pixel points in the minimum circumscribed rectangle are set to 1, and the pixel values of the pixel points outside the minimum circumscribed rectangle in the background template image are set to 0, to realize the mask processing of the background template image, and obtain a black-and-white template mask image, wherein the object contour region is a white rectangular region.
[0170] The advantage of this embodiment is that the object contour region of the reference object in the background template image is determined, the image position of the reference object in the background template image can be clearly determined, and the binarization processing (mask processing) of the background template image is realized by converting the pixel values of the pixel points, which can effectively expand the distinction degree of the pixel positions occupied by the reference object and the non-reference object in the background template image, so that the model can better explore the object contour features of the reference object based on the template mask image, thereby improving the accuracy of replacing the reference object with the target object.
[0171] The step 330 will be described in detail below.
[0172] In step 330, the denoising network control information of the image generation model is determined based on the contour features of the reference object in the background template image, the image description information, and the sample object image.
[0173] Please refer to FIG. 10, in one embodiment, the step 330 specifically includes but is not limited to the following steps 1010-1030:
[0174] Step 1010, encoding the image description information to obtain an image description embedding vector;
[0175] Step 1020, feature extraction is performed on the sample object image to obtain sample object feature data;
[0176] Step 1030, based on the image description embedding vector, the sample object feature data, and the contour feature of the reference object, control information is generated through a preset control network to obtain denoising network control information.
[0177] The steps 1010-1030 are described in detail below.
[0178] In step 1010, the image description embedding vector is used to indicate the vector representation of the conditional constraint of the denoising process in the image description information.
[0179] In the implementation of this embodiment, a text encoder can be used to encode the image description information, converting the image description information from data space to vector space to obtain the image description embedding vector. The text encoder includes a tokenizer, an embedding layer, and a text attention calculation module.
[0180] Specifically, first, the image description information is input into the tokenizer, and the tokenizer is used to tokenize the image description information to obtain a plurality of description words. Next, the embedding layer is used to perform word embedding processing on each description word, converting each description word into a vector form to obtain a description word vector corresponding to each description word. Finally, the text attention calculation module is used to perform attention calculation on the description word vectors of each description word to obtain the image description embedding vector.
[0181] In step 1020, the sample object feature data is used to indicate the object features possessed by the sample object; the sample object feature data can guide the denoising network of the image generation model to restore the object in the image to be more consistent with the sample object in the denoising process.
[0182] For example, when the sample object is a person, the sample object feature data includes but is not limited to indicating facial contours, facial features, and the like.
[0183] It should be noted that the sample object feature data can be obtained by a fine-tuning network generated by a fine-tuning technology (LoRA technology) based on a deep learning model. The fine-tuning network is a neural network based on cross-attention algorithm.
[0184] In the implementation of this embodiment, first, the sample object image is input into the fine-tuning network, and the fine-tuning network performs linear projection on the sample object image to generate a key vector, a value vector, and a query vector corresponding to the sample object image. Next, cross-attention calculation is performed using the key vector, the value vector, and the query vector to obtain a calculation result, and the calculation result is converted into a vector form to obtain the sample object feature data.
[0185] In step 1030, the preset control network is used to generate control information for the denoising network of the image generation model according to a plurality of input data. The input data allowed by the preset control network includes a conditional constraint of the image to be generated, an object sketch contour of the object in the background template image, an object feature of the object to be replaced, and the like.
[0186] It should be noted that, in order to improve the accuracy of the denoising network control information output by the control network, the control network of the embodiment of the present disclosure can also be a neural network based on cross-attention algorithm.
[0187] The denoising network control information is used to guide and control the attention and learning of the image generation model to each feature information in the denoising process, and the denoising network control information is also used to indicate the fine-tuning degree of various image feature information in the denoising process.
[0188] For the sake of brevity, the specific process of generating the denoising network control information by the preset control network will be described in detail hereinafter, and will not be described here.
[0189] The advantage of this embodiment is that the image description information, the contour feature of the reference object in the background template image, and the object feature of the sample object are used to generate the denoising network control information of the image generation model together, so that the denoising can be performed according to a plurality of constraint information in the image denoising process, and the accurate fine-tuning and flexible control of the image denoising are realized. This way can make the model perform image denoising under a plurality of dimensional conditional constraints, improve the image denoising ability of the model, be conducive to the model to generate image content more consistent with the image description, and also make the object in the image generated by the model more close to the action and posture of the reference object in the background template image.
[0190] Please refer to FIG. 11. In one embodiment, the contour feature of the reference object in the background template image is determined by the following way:
[0191] In step 1110, the object skeleton graph of the reference object is obtained by performing object detection on the background template image.
[0192] In step 1120, a plurality of object posture key points are obtained by performing posture feature extraction on the object skeleton graph.
[0193] In step 1130, the contour feature is determined based on the plurality of object posture key points.
[0194] The steps 1110-1130 will be described in detail below.
[0195] In step 1110, the object skeleton graph is used to indicate the skeleton contour of the reference object in the background template image.
[0196] In the implementation of the embodiment, first, an object is located in the background template image using a preset detection algorithm to identify the location of the reference object, and an object detection result is obtained, which is formed by a plurality of pixel points constituting the reference object. The preset detection algorithm includes but is not limited to a YOLO (You Only Look Once) algorithm, an SSD (Single Shot MultiBox Detector) algorithm, and the like. Then, the object skeleton graph of the reference object is determined according to the pixel points contained in the object detection result.
[0197] In step 1120, the object posture key point is used to indicate the action and posture of the reference object in the background template image.
[0198] In the implementation of the embodiment, first, the object skeleton graph is input into a preset posture estimation model, and the pixel points representing important parts of the reference object in the object skeleton graph are located in coordinates by using the preset posture estimation model, and a group of key point coordinates are obtained. Then, each key point coordinate output by the posture estimation model is determined as an object posture key point. The preset posture estimation model includes but is not limited to a PoseNet, an AlphaPose, and the like neural network model based on a deep learning algorithm.
[0199] In step 1130, first, the plurality of object posture key points are connected according to the relative positions and the correlation of the object posture key points, and an object line drawing graph is formed. Then, the object line drawing graph is converted into an image expression form acceptable to the control network, and the contour feature of the reference object is obtained.
[0200] As shown in FIG. 12, it is a key point detection process for a person (reference object) in a background template image. Specifically, the key point extraction algorithm is used to extract the key points of the person in the background template image, and the face key points and the skeleton key points of the reference object are obtained. The face key points are used to reflect the contour features and the facial features of the reference object, and the skeleton key points are used to reflect the posture features of the reference object in the background template image. Based on this, the face key points and the skeleton key points are jointly determined as the object posture key points of the reference object.
[0201] The advantage of the embodiment is that when the contour feature of the reference object is determined, the reference object is first located in the background template image, and the skeleton feature of the reference object is constructed according to the location of the reference object. Further, the key points are detected in the skeleton graph by using the posture estimation model, which can improve the determination efficiency and the determination accuracy of the posture key points, so that the contour feature of the reference object is drawn according to the relative positions of the plurality of posture key points, which can improve the determination accuracy of the contour feature, and is also beneficial to improve the accuracy of the control information of the denoising network.
[0202] Due to the fact that the image description information is converted into the form of embedding vector by using the conventional encoding mode, the condition constraint content of the generated image is often inaccurate. Embodiments of the present disclosure provide a scheme for encoding the image description information based on fine-tuning technology, which can improve the accuracy of the generated image description embedding vector, and further improve the effectiveness of the condition constraint content of the generated denoising network control information.
[0203] It should be noted that the encoding processing of the image description information in the embodiments of the present disclosure is implemented based on a preset text encoder, wherein the text encoder includes a tokenizer, an embedding layer and a text transformer.
[0204] Please refer to FIG. 13. In one embodiment, step 1010 specifically includes but is not limited to steps 1310-1340:
[0205] Step 1310, tokenizing the image description information to obtain a plurality of description words;
[0206] Step 1320, determining a target word in the plurality of description words, and finding a target word embedding feature corresponding to the target word based on a preset dictionary;
[0207] Step 1330, for each of the other description words in the plurality of description words except the target word, performing word embedding processing on each of the other description words to obtain a description word embedding feature of each of the other description words;
[0208] Step 1340, integrating the target word embedding feature and each of the description word embedding features into an image description embedding vector.
[0209] The steps 1310-1340 are described in detail below.
[0210] In step 1310, the description word is used to indicate the word information contained in the image description information.
[0211] In the implementation of this embodiment, first, the image description information can be input into the tokenizer of the text encoder. Then, the tokenizer will tokenize the image description information according to spaces, punctuation marks or delimiters to obtain a plurality of words. Further, the stop words in the tokenized plurality of words are removed to obtain a plurality of description words.
[0212] In step 1320, the target word refers to a word in the image description information that needs to be searched by a word table to find the corresponding embedding vector representation, and the target word embedding feature is used to indicate the vector representation form of the target word in the image description information.
[0213] The preset dictionary is preset in the embedding layer, and the preset dictionary of the embedding layer is used to indicate the correspondence between the index of each descriptive word and the word embedding feature.
[0214] In the implementation of this embodiment, first, for the plurality of descriptive words, the preset descriptive words to be represented by fixed characters are selected as target words from the plurality of descriptive words. Then, the index corresponding to the target word is input to the embedding layer of the text encoder, and the feature is searched in the preset dictionary of the embedding layer based on the index, and the word embedding feature corresponding to the index is found as the target word embedding feature.
[0215] In step 1330, the descriptive word embedding feature is used to indicate the vector representation of other descriptive words in the image description information.
[0216] In the implementation of this embodiment, for other descriptive words in the plurality of descriptive words except the target word, first, the index of each other descriptive word is determined according to the word segmentation result of the word segmenter. Then, the embedding layer of the text encoder is used to perform word embedding processing on each other descriptive word, and the word embedding feature corresponding to the index of each other descriptive word is searched in the preset dictionary of the embedding layer to obtain the descriptive word embedding feature of the other descriptive word.
[0217] In step 1340, first, the target word embedding feature and the descriptive word embedding feature are input to the text attention calculation module, and the descriptive word embedding vector of each descriptive word is output by the text attention calculation module. Then, according to the order of each descriptive word in the image description information, the number of each descriptive word is determined. Further, according to the number corresponding to the target word embedding feature and the number corresponding to the descriptive word embedding feature, the descriptive word embedding vectors are spliced in the order of the number to obtain a complete descriptive word embedding sequence. Finally, the spliced descriptive word embedding sequence is determined as the image description embedding vector.
[0218] As shown in FIG. 14, it is a specific illustration of the process of generating an image description vector as a conditional constraint based on image description information, and controlling image denoising based on the conditional constraint. Specifically, the input image instance is a clock image. The image description information of the input image instance is “A Photo of clock”, and “clock” in the image description information is taken as a target word, which is represented as “S * ”. At this time, the image description information becomes “A Photo of S *". Based on this, the image description information is input into the word segmenter, the image description information is segmented by the word segmenter, and the index of each word is determined according to the preset index corresponding to each candidate word in the preset vocabulary index table. Wherein, the index corresponding to the description word "A" is 508, the index corresponding to the description word "Photo" is 701, the index corresponding to the description word "of" is 73, the index corresponding to the description word "S * " is *. Further, the index of each description word is processed by the embedding layer to map the index corresponding to each description word from the numerical space to the vector space, to obtain the word embedding feature corresponding to each description word, wherein the word embedding feature corresponding to the description word "A" is "v 508 "; the word embedding feature corresponding to the description word "Photo" is "v 701 "; the word embedding feature corresponding to the description word "of" is "v 73 "; and the word embedding feature corresponding to the description word "S * " is "v * ". Further, the word embedding features of each description word are input into the text attention calculation module for attention calculation, and the image description vector c θ (y) is output by the text attention calculation module, the image description vector is used as a conditional constraint for denoising the noise image instance by the image generator, so that the predicted image instance output by the image generator is consistent with the content represented by the image description information. In addition, the predicted image instance is obtained by performing 4 diffusion processes (4 noise addition operations) on the input image instance to obtain a noise image instance, and then performing 4 denoising processes on the noise image instance by the image generator according to the image description vector.
[0219] The advantage of this embodiment is that the description word in the image description information which is pre-set to be represented by a fixed character is taken as a target word by the fine-tuning technology, and the word embedding representation of the target word is fine-tuned and searched by the text encoder to obtain the target word embedding feature corresponding to the target word, and the word embedding processing is performed on other description words to convert each description word into an embedding vector. This way can bind the object description content in the image description information to the image description embedding vector, which can improve the accuracy of the generated image description embedding vector, and further improve the effectiveness of the conditional constraint content used to generate the denoising network control information.
[0220] In the embodiments of the present disclosure, the control network includes a first control sub-network and a second control sub-network.
[0221] The first control sub-network and the second control sub-network are neural network structures that can perform cross-attention calculation. The network structures of the first control sub-network and the second control sub-network are the same, but the network parameters of the two are different.
[0222] The sample object feature data includes a first sample object feature, a second sample object feature, a third sample object feature, and a fourth sample object feature, wherein the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained by performing respective feature extraction on the sample object image.
[0223] The first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are all used for fine-tuning the denoising process of the image generation model, but the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature fine-tune different links. The first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are all obtained through the fine-tuning network described above. However, the network parameters of the fine-tuning network used for extracting the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are different.
[0224] Referring to FIG. 15, in one embodiment, step 1030 specifically includes but is not limited to steps 1510-1550:
[0225] Step 1510, inputting the image description embedding vector, the first sample object feature, and the contour feature into the first control sub-network to generate control information to obtain first control information;
[0226] Step 1520, inputting the image description embedding vector and the second sample object feature into the second control sub-network to generate control information to obtain second control information;
[0227] Step 1530, determining the up-sampling network control information based on the third sample object feature and the fourth sample object feature;
[0228] Step 1540, determining the down-sampling network control information based on the first control information, the second control information, and the first sample object feature and the second sample object feature;
[0229] Step 1550, integrating the up-sampling network control information and the down-sampling network control information into denoising network control information.
[0230] The steps 1510-1550 are described in detail below.
[0231] In step 1510, the first control information is used to fine-tune the intermediate result of the denoising network of the image generation model when down-sampling.
[0232] In the implementation of the embodiment, first, the image description embedding vector, the first sample object feature, and the contour feature are input to the first control subnetwork. Then, the first sample object feature and the contour feature are spliced to obtain a spliced feature vector. Further, the spliced feature vector is linearly projected by the first control subnetwork to obtain a query matrix vector; and the image description embedding vector is linearly projected by the first control subnetwork to obtain a key matrix vector and a value matrix vector. Further, cross-attention calculation is performed based on the query matrix vector, the key matrix vector, and the value matrix vector to obtain an attention calculation result. Finally, the first control information is obtained by converting the attention calculation result from a numerical space to a vector space through the first control subnetwork.
[0233] The specific process of performing cross-attention calculation based on the query matrix vector, the key matrix vector, and the value matrix vector to obtain the attention calculation result can be represented as shown in formula (1):
[0234] wherein Attention(Q, K, V) is the attention calculation result; Q is the query matrix vector; K is the key matrix vector, and V is the value matrix vector.d. k is the feature dimension of the key matrix vector, K T is the transposition result of the key matrix vector.
[0235] In step 1520, the second control information is used to fine-tune the result generated by the denoising network of the image generation model after downsampling.
[0236] In the implementation of the embodiment, the specific process of step 1520 is similar to step 1510 described above. The difference is that the input information of the first control subnetwork in step 1510 is the image description embedding vector, the first sample object feature, and the contour feature; while the input information of the second control subnetwork in step 1520 is the image description embedding vector and the second sample object feature, the input information of the two is different, and the control information generated by the two is different. In order to save space, no longer tedious.
[0237] In step 1530, the up-sampling network control information is used to conditionally constrain the up-sampling process of the denoising network of the image generation model, so as to fine-tune and correct the up-sampling process.
[0238] In the implementation of the embodiment, the third sample object feature and the fourth sample object feature are included in the same set, the object feature information in the set is regarded as a whole, and all the object feature information in the set is determined as the up-sampling network control information.
[0239] In step 1540, the down-sampling network control information is used to conditionally constrain the down-sampling process of the denoising network of the image generation model, so as to fine-tune and correct the down-sampling process.
[0240] In the implementation of this embodiment, the first control information, the second control information, the first sample object feature, and the second sample object feature are included in the same set, and all the information in the set is determined as the down-sampling network control information.
[0241] In step 1550, the up-sampling network control information and the down-sampling network control information are included in the same set, the control information in the set is regarded as a whole, and all the information in the set is determined as the denoising network control information.
[0242] As shown in FIG. 16, the sample object image is subjected to feature extraction by using the fine-tuning network with different model parameters, and the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained respectively. Then, the third sample object feature and the fourth sample object feature are integrated into the up-sampling network control information used to constrain the up-sampling process. Further, the first sample object feature, the second sample object feature, and the pre-acquired image description embedding vector and the contour feature of the reference object are input into the control network together, the first control information and the second control information are output by the control network, and the first control information, the second control information, the first sample object feature, and the second sample object feature are integrated into the down-sampling network control information used to constrain the down-sampling process. Based on this, the up-sampling network control information and the down-sampling network control information are used to realize the whole process control of the denoising processing of the denoising network, which can improve the accuracy of the denoising processing.
[0243] The advantage of this embodiment is that the denoising network control information is generated based on the image description information, the sample object feature in the sample object image, and the contour feature of the reference object in the background template image, so that the conditional constraint of the denoising network of the image generation model not only contains the description of the image to be generated, but also contains the object fine-tuning information determined according to the contour characteristics of the reference object and the object characteristics of the sample object, which is beneficial to improve the information comprehensiveness and information accuracy of the generated denoising network control information.
[0244] The step 340 is described in detail below.
[0245] In step 340, based on the sample splicing image feature and the denoising network control information, noise prediction is performed through the image generation model to obtain a noise prediction result of the noise reference image.
[0246] In the embodiments of the present disclosure, the image generation model includes a diffusion network and a denoising network.
[0247] The diffusion network is used to implement step-by-step noise adding processing on the sample splicing image feature, until the input sample splicing image feature approaches pure noise.
[0248] The denoising network is used to perform step-by-step denoising processing on the hidden space vector output by the diffusion network, until the image feature corresponding to the sample splicing image feature is generated.
[0249] Referring to FIG. 17, in one embodiment, step 340 specifically includes but is not limited to steps 1710-1730:
[0250] Step 1710, performing compression processing on the sample splicing image feature to obtain a sample compressed image feature;
[0251] Step 1720, performing diffusion processing on the sample compressed image feature based on the diffusion network to obtain a sample hidden space feature vector;
[0252] Step 1730, performing denoising processing on the sample hidden space feature vector through the denoising network based on the denoising network control information to obtain a noise prediction result.
[0253] The steps 1710-1730 are described in detail below.
[0254] In step 1710, since the sample splicing image feature is spliced from three image features, the feature dimension of the sample splicing image feature is higher than the feature dimension of the input feature that the diffusion network can accept. Based on this, first, the sample splicing image feature is compressed to reduce the feature dimension of the sample splicing image feature, without changing the image feature information possessed by the sample splicing image feature, to obtain a sample compressed image feature. The feature dimension of the sample compressed image feature is W*H*4, which meets the feature dimension of the input feature that the diffusion network can accept.
[0255] In step 1720, the sample hidden space feature vector is used to indicate the result of adding noise to the sample splicing image feature by the diffusion network at a fixed time step.
[0256] In the implementation of this embodiment, when the diffusion network performs diffusion processing on the sample compressed image feature, the diffusion network adds noise to the sample compressed image feature step by step until the sample compressed image feature approaches pure noise. The diffusion process of the embodiment of the present disclosure as a whole can be a parameterized Markov chain.
[0257] For example, the fixed time step is set as T, the sample compressed image feature is added with noise T times through the forward process of the diffusion network to generate the hidden space representation corresponding to the noise reference image, and the generated hidden space representation is determined as the sample hidden space feature vector, where T is a positive integer.
[0258] Specifically, noise is added to the sample compressed image feature one by one through the diffusion process of the diffusion network, and the sample compressed image feature loses its features one by one. After T times of adding noise, the sample compressed image feature becomes a hidden space representation without any features. The hidden space representation is determined as the sample hidden space feature vector.
[0259] It should be noted that the sample hidden space feature vector refers to the representation of a pure noise image without image features corresponding to the noise reference image. The form of the sample hidden space feature vector is the same as that of the object representation, which can be a vector representation or a matrix representation, and is not limited.
[0260] In step 1730, when the sample hidden space feature vector is de-noised by the de-noising network, the de-noising network control information is used as a constraint condition to remove noise from the sample hidden space feature vector one by one until the sample hidden space feature vector is restored to an image feature that meets the constraint requirements of the image description information. The backward process of the de-noising network as a whole can also be a parameterized Markov chain.
[0261] To save space, the specific process of de-noising the sample hidden space feature vector by the de-noising network to obtain the noise prediction result in the embodiment of the present disclosure will be described in detail below. Herein, no further description is given.
[0262] As shown in FIG. 18, it is a noise prediction process of the image generation model. Specifically, the sample spliced image feature is input into the image generation model. First, the image generation model compresses the sample spliced image feature to form a sample compressed image feature with a smaller feature dimension than that of the sample spliced image feature. Then, the sample compressed image feature is added with noise T times (i.e., diffusion process) through the diffusion network to convert the sample compressed image feature into an image feature close to pure noise, thereby obtaining a sample hidden space feature vector. The sample hidden space feature vector cannot reflect the image information in the sample spliced image feature. Further, the sample hidden space feature vector is de-noised T times (i.e., de-noising process) through the de-noising network based on the given de-noising network control information to eliminate the noise carried on the sample hidden space feature vector, thereby restoring the sample hidden space feature vector to an image feature form and obtaining a noise prediction result. The noise prediction result can reflect the image information in the sample spliced image feature to a certain extent.
[0263] The embodiment has the advantages that the sample splicing image features are compressed into sample compressed image features first, so that the input sample compressed image features meet the input requirements of the image generation model. Then, the diffusion network is used to add noise to the sample compressed image features multiple times, so as to reduce the details and clarity of the sample compressed image features, and obtain the sample latent space feature vector. Further, the denoising network control information is used as a constraint condition, and the sample latent space feature vector is sequentially removed noise by the denoising network until the sample latent space feature vector is restored to the image features meeting the constraint requirements of the image description information, and the noise prediction result is obtained. This way introduces the image description information, the contour features of the reference object and the object features of the sample object to fine-tune the denoising process, which can effectively improve the denoising ability of the model and is beneficial to improve the accuracy of the target image generated by the model.
[0264] In the embodiments of the present disclosure, the denoising network can be a U-shaped network structure (U-net network), and the denoising network includes an up-sampling attention network and a down-sampling attention network. The denoising network control information includes up-sampling network control information and down-sampling network control information.
[0265] The down-sampling attention network is used for down-sampling processing of the sample latent space feature vector to obtain more low-dimensional features.
[0266] The up-sampling attention network is used for up-sampling processing of the output result of the up-sampling attention network to restore the output result to the denoising image features with the same feature dimension as the sample latent space feature vector.
[0267] Please refer to FIG. 19. In one embodiment, step 1730 specifically includes but is not limited to steps 1910-1920:
[0268] Step 1910, the down-sampling network control information is fused into the first attention matrix of the down-sampling attention network to update the first attention matrix, and the up-sampling network control information is fused into the second attention matrix of the up-sampling attention network to update the second attention matrix.
[0269] Step 1920, the sample latent space feature vector is denoised by the down-sampling attention network after updating the first attention matrix and the up-sampling attention network after updating the second attention matrix, and the noise prediction result is obtained.
[0270] The steps 1910-1920 are described in detail below.
[0271] In step 1910, the first attention matrix is used for attention calculation in the down-sampling process of the denoising network to capture the degree of correlation between different input features (sample latent space feature vectors, image description embedding vectors); the second attention matrix is used for attention calculation in the up-sampling process of the denoising network to capture the degree of correlation between different input features (output results of the down-sampling attention network, image description embedding vectors).
[0272] In the implementation of this embodiment, the down-sampling attention network includes an attention down-sampling module, and the attention down-sampling module includes a first attention matrix and a residual block structure. When the first attention matrix is updated, the first sample object feature and the second sample object feature in the down-sampling network control information are weighted to the output of the first attention matrix, and the first control information and the second control information in the down-sampling network control information are weighted to the output of the attention down-sampling module. Similarly, the up-sampling attention network includes an attention up-sampling module, and the attention up-sampling module includes a second attention matrix and a residual block structure. When the second attention matrix is updated, the third sample object feature and the fourth sample object feature in the up-sampling network control information are weighted to the output of the second attention matrix.
[0273] As shown in FIG. 20, the down-sampling attention network includes two attention down-sampling modules, and each of the two attention down-sampling modules has a first attention matrix, the first sample object feature is weighted to the output of the first first attention matrix, and the second sample object feature is weighted to the output of the second first attention matrix. At the same time, the first control information is weighted to the output of the first attention down-sampling module, and the second control information is weighted to the output of the second attention down-sampling module.
[0274] As shown in FIG. 21, the up-sampling attention network includes two attention up-sampling modules, and each of the two attention up-sampling modules has a second attention matrix, the third sample object feature is weighted to the output of the first second attention matrix, and the fourth sample object feature is weighted to the output of the second second attention matrix.
[0275] In step 1920, when the sample latent space feature vector is denoised by the down-sampling attention network after the first attention matrix is updated and the up-sampling attention network after the second attention matrix is updated, first, the sample latent space feature vector and the image description embedding vector are input into the denoising network of the image generation model, the sample latent space feature vector is linearly projected to obtain the query feature Q corresponding to the sample latent space feature vector, and the image description embedding vector is linearly projected to obtain the key feature K and the value feature V corresponding to the image description embedding vector. Then, the first attention matrix of the first attention down-sampling module of the down-sampling attention network is used to perform cross-attention calculation on the query feature Q, the key feature K and the value feature V to obtain a first attention calculation result, and the first attention calculation result and the first sample object feature are weighted according to a preset weight ratio to obtain a first weighted result. Further, the first weighted result is processed by the residual block structure of the first attention down-sampling module to obtain a first residual processing result, and the first control information and the first residual processing result are weighted according to a preset weight to obtain a second weighted result. Then, the second weighted result is linearly projected to obtain a new query feature Q, and the first attention matrix of the second attention down-sampling module of the down-sampling attention network is used to perform cross-attention calculation on the key feature K, the value feature V and the new query feature Q to obtain a second attention calculation result, and the second attention calculation result and the second sample object feature are weighted according to a preset weight ratio to obtain a third weighted result. Further, the third weighted result is processed by the residual block structure of the second attention down-sampling module to obtain a second residual processing result, and the second control information and the second residual processing result are weighted according to a preset weight to obtain a fourth weighted result, wherein the fourth weighted result is the output result of the down-sampling attention network.
[0276] Further, the fourth weighted result and the image description embedding vector are input into the up-sampling attention network, the fourth weighted result is first linearly projected to obtain a query feature, and the image description embedding vector is linearly projected to obtain a key feature and a value feature corresponding to the image description embedding vector. Then, the key feature, the value feature and the query feature are cross-attention calculated by using a second attention matrix of a first attention up-sampling module of the up-sampling attention network to obtain a third attention calculation result, and the third attention calculation result and the third sample object feature are weighted calculated according to a preset weight ratio to obtain a fifth weighted result. Further, the fifth weighted result is residual processed by a residual block structure of the first attention up-sampling module to obtain a third residual processing result. Finally, the third residual processing result is linearly projected to obtain a new query feature, and the key feature, the value feature and the new query feature are cross-attention calculated by using a second attention matrix of a second attention up-sampling module of the up-sampling attention network to obtain a fourth attention calculation result, and the fourth attention calculation result and the fourth sample object feature are weighted calculated according to a preset weight ratio to obtain a sixth weighted result. Further, the sixth weighted result is residual processed by a residual block structure of the second attention up-sampling module to obtain a denoising result of the sample hidden space feature vector generated by the denoising network at the prediction time step T; the above process is repeated, and the denoising result of the sample hidden space feature vector is further denoised for T-1 times to obtain a noise prediction result.
[0277] The embodiment has the advantage that corresponding denoising network control information is introduced at different links of the denoising process to fine-tune the denoising process. Specifically, the first sample object feature and the second sample object feature in the down-sampling network control information are fused into the first attention matrix of the down-sampling attention network, and the third sample object feature and the fourth sample object feature in the up-sampling network control information are fused into the second attention matrix of the up-sampling attention network, so that the output results of multiple attention matrices in the denoising network can be corrected. In addition, the output results of the down-sampling attention network are corrected by the first control information and the second control information, and multiple fine-tuning is performed in each denoising process, so that the stability and accuracy of noise prediction can be improved, the noise prediction result finally generated can be more in line with the actual demand, and the model training effect can be improved.
[0278] Step 350 is described in detail below.
[0279] In step 350, the image generation model is trained based on comparison of the noise reference images and the noise prediction results of the plurality of image-text sample pairs.
[0280] Referring to FIG. 22, in one embodiment, step 350 specifically includes but is not limited to steps 2210-2230.
[0281] Step 2210, for each image-text sample pair, reference noise in the noise reference image thereof is obtained, and based on a comparison result of the reference noise and a noise prediction result corresponding to the noise reference image, a sub-loss function of the image-text sample pair is calculated.
[0282] Step 2220, based on the sub-loss function of each image-text sample pair, a total loss function is determined.
[0283] Step 2230, based on the total loss function, the image generation model is trained.
[0284] Steps 2210-2230 are described in detail as follows.
[0285] In step 2210, the reference noise is used to indicate a degree of adding noise to the reference image, and the sub-loss function is used to indicate a difference degree between the reference noise and the predicted noise of a single image-text sample pair.
[0286] In the implementation of this embodiment, first, for each image-text sample pair, noise extraction is performed on the noise reference image thereof to obtain the reference noise in the noise reference image. Then, noise difference calculation is performed on the reference noise and the noise prediction result to obtain a noise difference result. Finally, the sub-loss function is calculated according to the noise difference result.
[0287] In step 2220, the total loss function is used to indicate an overall difference degree between the reference noise and the predicted noise of all image-text sample pairs. The smaller the total loss function is, the smaller the overall difference between the reference noise and the predicted noise of all image-text sample pairs is, and the higher the image generation accuracy of the image generation model is.
[0288] In the implementation of this embodiment, the sub-loss functions of all image-text sample pairs are averaged to obtain the total loss function. Specifically, first, the total number of image-text sample pairs is determined. Then, the sub-loss functions of all image-text sample pairs are added to obtain a sum of the sub-loss functions. Finally, the sum of the sub-loss functions is divided by the total number of image-text sample pairs to obtain the total loss function.
[0289] In step 2230, the model parameters of the image generation model are adjusted to minimize the total loss function, and steps 310-350 are repeated to realize iterative training of the image generation model. The model parameters that minimize the total loss function are taken as the final model parameters, and the image generation model with the final model parameters is taken as the trained image generation model.
[0290] The embodiment has the advantages that, based on the manner of supervised learning, the sub-loss function of each image-text sample pair is determined according to the noise difference between the reference noise and the noise prediction result of each image-text sample pair, and the total loss function is constructed based on the plurality of sub-loss functions, which can train the image generation model by minimizing the difference between the reference noise and the noise prediction result, and is beneficial to improve the model training effect, and further improve the prediction accuracy of the model on the generated image.
[0291] In the embodiment of the present disclosure, the noise prediction result is obtained through the prediction of the plurality of prediction time steps.
[0292] The prediction time step is used to indicate the noise adding frequency and the noise removing frequency of the sample splicing image feature of the input image generation model.
[0293] Please refer to FIG. 23, in one embodiment, step 2210 specifically includes but is not limited to steps 2310-2330:
[0294] Step 2310, determining the prediction noise of the last prediction time step based on the noise prediction result;
[0295] Step 2320, performing regular term calculation based on the reference noise and the prediction noise to obtain a regular term calculation result;
[0296] Step 2330, determining the sub-loss function based on the regular term calculation result.
[0297] The steps 2310-2330 are described in detail as follows.
[0298] In step 2310, the prediction noise is used to indicate the noise contained in the denoising result (denoising image) generated by the image generation model at the last prediction time step.
[0299] In the implementation of the embodiment, since the image generation model will denoise the denoising result (denoising image) generated at the previous prediction time step again at each prediction time step, and the noise prediction result is obtained through the denoising processing of the plurality of prediction time steps, the noise contained in the denoising result of different prediction time steps is different. Based on this, the noise prediction result contains the denoising result (denoising image) generated at the last prediction time step. The prediction noise of the last prediction time step can be directly extracted from the noise prediction result.
[0300] In step 2320, the regular term calculation result is used to indicate the difference degree between the reference noise and the prediction noise.
[0301] In the implementation of this embodiment, first, the reference noise and the predicted noise are subtracted to obtain a noise difference for each text-image sample pair. Then, the noise difference is regularized to obtain a regularization term calculation result.
[0302] In step 2330, the regularization term calculation result is converted based on the denoising network control information, the regularization term calculation result is converted into the form of a conditional loss function, and the converted loss function is determined as a sub-loss function. The sub-loss function of the embodiment of the present disclosure can be represented as formula (2):
[0303] Wherein, L LDM is a sub-loss function; ε is a reference noise, represents that the reference noise is subject to a standard normal distribution (mean is 0, variance is 1). t represents a prediction time step, which is used to update the sample latent space feature vector step by step in the image generation process. z t represents the sample latent space feature vector at the prediction time step t, z t is the result of adding noise to the sample compressed image feature z at the prediction time step t. y and c both refer to denoising network control information, which is used to conditionally constrain the image generation process. x is a noise reference image used for training. is used to indicate the expectation of the encoder ε for the input x. θ refers to the predicted noise after the sample latent space feature vector z t is denoised according to the conditional constraint c at the prediction time step t. is a regularization term calculation result.
[0304] The advantage of this embodiment is that by calculating the regularization term loss between the reference noise and the predicted noise of the text-image sample pair, obtaining the regularization term calculation result of each text-image sample pair, using the regularization term calculation result (L2 loss) as a sub-loss function to train the image generation model, the denoising network control information can be used to better adjust the image generation process of the image generation model, so that the generation process is more controllable. At the same time, the conditional information (denoising network control information) is added to the sub-loss function, which can make the predicted noise closely related to the given condition, improve the consistency and accuracy of the generated image, and further improve the image quality of the target image generated by the model.
[0305] As shown in FIG. 24, it is a whole flow chart of model training of the embodiment of the present disclosure. Specifically, the background template image, the expected effect image, and the image description information of the expected effect image are taken as inputs.
[0306] First, the reference noise is generated using a random number i, and the reference noise is added to the intended effect diagram to obtain a noise reference image, the specific process is similar to the above steps 510-530. Next, the background template image is masked to obtain an object replacement mask (template mask image), the specific process is similar to the above steps 810-820.
[0307] Further, the template mask image, the noise reference image, and the background template image are respectively encoded by the encoder, and the three encoding results are spliced into a sample splicing image feature. The sample splicing image feature is compressed into a sample compressed image feature Z by the image generation model, and the sample compressed image feature Z is diffused by the diffusion network. The sample compressed image feature Z is added with noise T times to obtain a sample latent space feature vector Z T . Further, the image description information of the intended effect diagram is converted into a text form conforming to the input requirements of the text encoder τ, and the image description information of the intended effect diagram is input into the text encoder. The image description embedding vector is output by the text encoder, the specific process is similar to the above steps 1310-1340. Further, the object box corresponding to the background template image is extracted, and the contour line drawing of the reference object in the background template image (contour feature) is generated based on the extracted object box. The specific process is similar to the above steps 1110-1130. At the same time, based on the sample object image, sample object feature data is generated, wherein the sample object feature data includes first object feature QKV1-A*, second object feature QKV2-A*, third object feature QKV3-A* and fourth object feature QKV4-A*. Further, the first object feature QKV1-A*, the second object feature QKV2-A*, the contour feature and the image description embedding vector are input into the control network, and the first control information is output by the first control subnetwork of the control network, and the second control information is output by the second control subnetwork of the control network.
[0308] Next, when the denoising network denoises the sample latent space feature vector Z T , first, based on the first control information, the first object feature QKV1-A*, the second object feature QKV2-A*, the sample latent space feature vector Z T is down-sampled by each attention matrix of the down-sampling attention network of the denoising network to obtain a down-sampling result, and the down-sampling result is modified using the second control information to obtain a modified down-sampling result; then, based on the third object feature QKV3-A* and the fourth object feature QKV4-A*, the modified down-sampling result is up-sampled by each attention matrix of the up-sampling attention network of the denoising network to obtain a denoising result Z T-1' at the prediction time step T. Further, the denoising result ZT-1' After T-1 times of denoising, the denoising result Z is obtained according to the above process ’ The specific process is similar to steps 1910-1920.
[0309] Finally, based on the denoising result Z ’ noise prediction is performed to obtain a noise prediction result, a loss function LOSS is constructed according to the reference noise and the noise prediction result, and the image generation model is iteratively trained based on the loss function LOSS until the image generation model meets the training requirements. The specific process is similar to step 350. For the sake of brevity, no further description is given.
[0310] The image generation method of one embodiment of the present disclosure is described in detail below.
[0311] According to one embodiment of the present disclosure, an image generation method is provided.
[0312] The image generation method is generally applied in a business scenario in which an object in a fixed background image is to be replaced by a target person, a target object, or the like, such as a video production scenario, an object display scenario, and the like, as shown in FIGS. 2A-2C. The present embodiment provides a scheme for generating an image based on image description information and object information (contour information of a reference object in a background image and object features of a target object to be replaced) through an image generation model, which can improve the accuracy of target image generation.
[0313] As shown in FIG. 25, the image generation method according to one embodiment of the present disclosure can be executed by an electronic device, which can be an image processing server or an object terminal shown in FIG. 1, and can include the following steps:
[0314] Step 2510, obtaining an object image including a first object, a background image, and description information;
[0315] Step 2520, determining a spliced image feature based on the background image, a preset noise image, and a mask image;
[0316] Step 2530, determining denoising control information of an image generation model based on contour features of a second object in the background image, the description information, and the object image;
[0317] Step 2540, generating an image through the image generation model based on the spliced image feature and the denoising control information to obtain a target image.
[0318] The steps 2510-2540 are described in detail below.
[0319] In step 2510, an object image including the first object, a background image, and description information are acquired.
[0320] The object image including the first object refers to an image that can reflect the object characteristics of the first object, and one or more object images including the first object can be acquired.
[0321] The background image refers to an image that provides background content for the final image to be generated, that is, the background image to which the first object is to be added, wherein the background image contains a second object to be replaced by the first object in the object image.
[0322] The description information is used to describe the replacement from the second object to the first object.
[0323] In the implementation of this embodiment, the specific process of step 2510 is similar to the specific process of acquiring the sample object image including the sample object, the background template image, and the image description information in step 310 described above. To save space, no longer tedious.
[0324] In step 2520, the splicing image feature is determined based on the background image, the preset noise image, and the mask image.
[0325] The preset noise image is a noise image generated based on random numbers and subject to Gaussian distribution, and its specific generation method is similar to that of step 520 described above. The difference is that the random numbers used to generate the preset noise image in step 2520 are different from those in step 520. To save space, no longer tedious.
[0326] The mask image is obtained by masking the second object in the background image, and its specific process is similar to steps 810-820 described above. To save space, no longer tedious.
[0327] The splicing image feature is used to indicate the feature splicing result of the background image, the preset noise image, and the mask image.
[0328] To save space, the specific process of determining the target splicing image feature based on the background image, the preset noise image, and the mask image of the embodiment of the present disclosure will be described in detail below. No longer tedious here.
[0329] In step 2530, the denoising control information of the image generation model is determined based on the contour feature of the second object in the background image, the description information, and the object image.
[0330] The image generation model is generated according to the training method of the image generation model of the above-mentioned embodiment.
[0331] The denoising control information is used as a conditional constraint to assist the image generation model in image denoising when generating an image, so as to improve the image denoising accuracy and make the image denoising effect meet the actual requirements.
[0332] In the implementation of this embodiment, the specific process of step 2530 is similar to that of step 330 described above. For brevity, no longer be described.
[0333] In step 2540, based on the spliced image feature and the denoising control information, an image generation model is used to generate an image to obtain a target image.
[0334] The target image is used to indicate the result of replacing the second object in the background image with the first object of the object image.
[0335] For brevity, the specific process of the embodiment of the disclosure based on the spliced image feature and the denoising control information, and the image generation model is used to generate an image to obtain a target image will be described in detail below. Here, no longer be described.
[0336] Through steps 2510-2540, in the embodiment of the disclosure, when the image generation model is used to generate an image, the image features of the background image, the preset noise image, and the mask image are integrated into the spliced image feature, so that the background image information and the second object information of the background image in the spliced image feature are fused, and the noise information is also fused in the spliced image feature, which can make the spliced image feature meet the actual situation. Further, the contour feature of the second object in the background image, the description information, and the first object image are introduced to generate the denoising control information of the denoising network of the image generation model, which can make the image generation model adjust the object in the denoising link, so that the object in the generated target image meets the conditional constraints of the description information, the object characteristics of the first object, and the contour characteristics of the second object, so that the first object and the background of the target image have good consistency and coordination, thereby improving the accuracy of the model in generating the target image.
[0337] Please refer to FIG. 26, in one embodiment, step 2520 specifically includes but is not limited to the following steps 2610-2640:
[0338] Step 2610, performing first encoding processing on the preset noise image to obtain noise image features;
[0339] Step 2620, performing second encoding processing on the background image to obtain background image features;
[0340] Step 2630, performing third encoding processing on the mask image to obtain mask image features;
[0341] Step 2640, splicing the noise image feature, the background image feature, and the mask image feature to obtain a spliced image feature.
[0342] The steps 2610-2640 are described in detail below.
[0343] The noise image feature is used to indicate a result of converting the preset noise image from the data space to the pixel space.
[0344] The background image feature is used to indicate a result of converting the background image from the data space to the pixel space.
[0345] The mask image feature is used to indicate a result of converting the mask image from the data space to the pixel space.
[0346] In the implementation of this embodiment, the specific processes of the steps 2610-2640 are similar to those of the steps 710-740 described above. For brevity, no longer be described.
[0347] The advantage of this embodiment is that the image information of the preset noise image, the background image, and the mask image is converted from the pixel space to the latent vector space, and the image information (the noise image feature, the background image feature, and the mask image feature) of the preset noise image, the background image, and the mask image in the latent vector space is spliced to form a spliced image feature with multiple image information. Further, the spliced image feature is used as the input of the image generation model during image generation, so that the noise information, the background information, and the second object information in the background are fused in the input data of the model, which can better improve the richness and comprehensiveness of the feature information of the spliced image feature, and is conducive to improving the image generation accuracy of the model, so that the finally generated target image is more real and more accurate.
[0348] In the embodiments of the present disclosure, the image generation model includes a diffusion network, a denoising network, and a decoding network.
[0349] The decoding network is used to convert the image feature in the denoising result from the latent vector space to the pixel space.
[0350] Please refer to FIG. 27. In one embodiment, the step 2540 specifically includes but is not limited to the following steps 27710-2740:
[0351] Step 2710, compressing the spliced image feature to obtain a compressed image feature;
[0352] Step 2720, diffusing the compressed image feature based on the diffusion network to obtain a latent space feature vector;
[0353] Step 2730, denoising the latent space feature vector based on the denoising control information through the denoising network to obtain a denoising result;
[0354] Step 2740, feature decoding the denoising result based on the decoding network to obtain the target image.
[0355] The steps 2710-2740 are described in detail below.
[0356] The compressed image feature is used to indicate the feature dimension reduction result of the spliced image feature. The compressed image feature has the same image feature information as the spliced image feature, but the feature dimension of the compressed image feature is lower than that of the spliced image feature.
[0357] The latent space feature vector is used to indicate the result of adding noise to the compressed image feature at a fixed time step by the diffusion network.
[0358] The denoising result is used to indicate the image feature that satisfies the constraint requirement of the image description information generated by denoising the latent space feature vector at a fixed time step by the denoising network.
[0359] In the implementation of this embodiment, the specific process of steps 2710-2730 is similar to that of steps 1710-1730 described above. To save space, it will not be described again.
[0360] In step 2740, the denoising result is input to the decoding network, and the image feature in the denoising result is mapped back to the original pixel space by the decoding network to generate a noise-free target image consistent with the description information.
[0361] The advantage of this embodiment is that the spliced image feature is first compressed into a compressed image feature, so that the input compressed image feature meets the input requirement of the image generation model. Then, the diffusion network is used to add noise to the compressed image feature multiple times to reduce the details and clarity of the compressed image feature to obtain a latent space feature vector. Further, the denoising network is used to denoise the latent space feature vector multiple times to predict the source image. In this process, the denoising control information is used to fine-tune the predicted image feature, so that the object feature in the final predicted denoising result is more consistent with the true requirement. Finally, the decoding network is used to decode the denoising result to obtain the predicted target image, which can improve the image quality of the target image and make the objects and backgrounds in the target image have better coordination.
[0362] As shown in FIG. 28, it is a specific application module diagram of the image generation method of the embodiment of the present disclosure. Specifically, first, for a certain background image containing a second object, the background image is subjected to object face mask processing to obtain a mask image, and the specific process is similar to steps 810-820 described above; the background image is subjected to object mask part key points to obtain the contour features of the second object of the background image, and the specific process is similar to steps 1110-1130 described above. In addition, for the first object to be replaced with the second object, the object image including the first object is subjected to feature extraction to obtain object fine-tuning feature data of the first object. Further, based on the random noise image, the background image and the mask image, the latent space input is constructed to obtain the splicing image feature, and the specific process is similar to steps 2610-2640 described above. Then, the contour features, the description information of the target image to be generated, etc. are used to drive the control network, and the denoising control information output by the control network is used as auxiliary (conditional constraint), and the splicing image feature is subjected to diffusion processing and denoising processing by the image generation model to obtain a denoising result. Finally, the denoising result is converted into a target image by the decoding network for output, so as to generate a target image with the background image as the background and the first object as the foreground, and the specific process is similar to steps 2710-2740 described above. For the sake of brevity, no longer repeated.
[0363] As shown in FIG. 29, it is a whole flow chart of the image generation based on the image generation model of the embodiment of the present disclosure. Specifically, the background image, and the description information corresponding to the image to be generated are taken as inputs.
[0364] First, a random noise image is generated by using a random number i. Then, the background image is masked to obtain an object replacement mask (mask image), and the specific process is similar to steps 810-820 described above. Further, the random noise image, the mask image, and the background image are respectively encoded by using the encoder, and the three encoding results are spliced into a splicing image feature.
[0365] Further, the splicing image feature is compressed into a compressed image feature Z by the image generation model, and the diffusion network is used to diffuse the compressed image feature Z, and the compressed image feature Z is added with noise for T times to obtain a latent space feature vector Z TFurther, the description information corresponding to the image to be generated is taken as a conditional constraint, converted into a text form conforming to the input requirements of the text encoder τ, and input to the text encoder, and the image description embedding vector is output by the text encoder. The specific process is similar to steps 1310-1340 described above. Further, the object box corresponding to the background image is extracted, and the contour sketch graph (contour feature) of the second object in the background image is generated based on the extracted object box. The specific process is similar to steps 1110-1130 described above. At the same time, based on the object image including the first object, the object feature data is generated, wherein the object feature data includes the first object feature QKV1-A*, the second object feature QKV2-A*, the third object feature QKV3-A* and the fourth object feature QKV4-A*. Further, the first object feature QKV1-A*, the second object feature QKV2-A*, the contour feature and the image description embedding vector are input to the control network, and the first control information is output by the first control subnetwork of the control network, and the second control information is output by the second control subnetwork of the control network.
[0366] Next, when the denoising network denoises the latent space feature vector Z T , the first control information, the first object feature QKV1-A* and the second object feature QKV2-A* are used to perform downsampling on the latent space feature vector Z T through each attention matrix of the downsampling attention network of the denoising network, to obtain a downsampling result, and the second control information is used to correct the downsampling result to obtain a corrected downsampling result; and then the third object feature QKV3-A* and the fourth object feature QKV4-A* are used to perform upsampling on the corrected downsampling result through each attention matrix of the upsampling attention network of the denoising network, to obtain the denoising result Z T-1' at the prediction time step T. T-1' Further, the denoising result Z ’ at the prediction time step T is denoised according to the above process, and after T-1 times of denoising, the denoising result Z ’ is obtained. The specific process is similar to steps 1910-1920 described above. Finally, the denoising result Z ’ is feature decoded based on the decoding network to obtain the target image I. The specific process is similar to step 2740 described above. For the sake of brevity, no longer described.
[0367] The apparatus and device of the embodiment of the present disclosure are described below.
[0368] It can be understood that, although each step in each of the above flowcharts is shown in sequence according to the representation of the arrow, these steps are not necessarily executed in the order represented by the arrow. Unless otherwise specified in the embodiments, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least part of the steps in the above flowcharts can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or steps or stages in other steps.
[0369] It should be noted that, in each specific embodiment of the present application, when it is necessary to perform relevant processing according to data related to the characteristics of the target object, such as target object attribute information or attribute information set, the permission or consent of the target object is obtained first, and the collection, use and processing of the data comply with relevant laws, regulations and standards. In addition, when the embodiments of the present application need to obtain target object attribute information, the separate permission or separate consent of the target object is obtained through a pop-up window or by jumping to a confirmation page, and after obtaining the separate permission or separate consent of the target object, the necessary target object related data for enabling the embodiments of the present application to operate normally is obtained.
[0370] FIG. 30 is a structural schematic diagram of a training device 3000 of an image generation model provided by the embodiments of the present disclosure. The training device 3000 of the image generation model comprises:
[0371] A first acquisition unit 3010 is configured to acquire a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image comprising a sample object, and the noise reference image is obtained by adding noise to a reference object obtained by replacing the sample object in the background template image.
[0372] A first determination unit 3020 is configured to determine sample splicing image features based on the noise reference image, the background template image, and a template mask image, wherein the template mask image is obtained by masking the reference object in the background template image.
[0373] A second determination unit 3030 is configured to determine denoising network control information of the image generation model based on contour features of the reference object in the background template image, the image description information, and the sample object image.
[0374] The prediction unit 3040 is configured to perform noise prediction on the noise reference image based on the sample spliced image feature and the denoising network control information, to obtain a noise prediction result corresponding to the noise reference image.
[0375] The training unit 3050 is configured to train the image generation model based on comparison results of the noise reference images and the respective noise prediction results in the plurality of image-text sample pairs.
[0376] Optionally, the training unit 3050 comprises:
[0377] The calculation module (not shown) is configured to, for each image-text sample pair, obtain reference noise in the noise reference image in the image-text sample pair, and calculate a sub-loss function of the image-text sample pair based on a comparison result of the reference noise and a noise prediction result corresponding to the noise reference image in the image-text sample pair, the reference noise being noise added in the reference image.
[0378] The determination module (not shown) is configured to determine a total loss function based on the sub-loss function of each image-text sample pair.
[0379] The training module (not shown) is configured to train the image generation model based on the total loss function.
[0380] Optionally, the noise prediction result is obtained through prediction of a plurality of prediction time steps.
[0381] The calculation module (not shown) is configured to:
[0382] Determine prediction noise of the last prediction time step based on the noise prediction result.
[0383] Perform regular term calculation based on the reference noise and the prediction noise to obtain a regular term calculation result.
[0384] Determine the sub-loss function based on the regular term calculation result.
[0385] Optionally, the first determination unit 3020 is configured to:
[0386] Perform first encoding processing on the noise reference image to obtain reference image encoding features.
[0387] Perform second encoding processing on the background template image to obtain background image encoding features.
[0388] Perform third encoding processing on the template mask image to obtain mask image encoding features.
[0389] Splice the reference image encoding features, the background image encoding features, and the mask image encoding features to obtain the sample spliced image feature.
[0390] Optionally, the second determination unit 3030 comprises:
[0391] An encoding module (not shown) is configured to encode the image description information to obtain an image description embedding vector;
[0392] An extraction module (not shown) is configured to extract features of the sample object image to obtain sample object feature data;
[0393] A generation module (not shown) is configured to generate control information by using a preset control network based on the image description embedding vector, the sample object feature data, and the contour feature of the reference object, to obtain denoising network control information.
[0394] Optionally, the encoding module (not shown) is configured to:
[0395] Tokenize the image description information to obtain a plurality of description words;
[0396] Determine a target word from the plurality of description words, and find a target word embedding feature corresponding to the target word based on a preset dictionary;
[0397] For each other description word in the plurality of description words except the target word, perform word embedding processing on each other description word to obtain a description word embedding feature of each other description word;
[0398] Integrate the target word embedding feature and the description word embedding feature into the image description embedding vector.
[0399] Optionally, the control network includes a first control sub-network and a second control sub-network; and the sample object feature data includes a first sample object feature, a second sample object feature, a third sample object feature, and a fourth sample object feature, wherein the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained by performing respective feature extraction on the sample object image.
[0400] The generation module (not shown) is configured to:
[0401] Input the image description embedding vector, the first sample object feature, and the contour feature into the first control sub-network to generate control information, to obtain first control information;
[0402] Input the image description embedding vector and the second sample object feature into the second control sub-network to generate control information, to obtain second control information;
[0403] Determine up-sampling network control information based on the third sample object feature and the fourth sample object feature;
[0404] Determine down-sampling network control information based on the first control information, the second control information, and the first sample object feature and the second sample object feature;
[0405] The up-sampling network control information and the down-sampling network control information are integrated into the denoising network control information.
[0406] Optionally, the noise reference image is generated by:
[0407] A reference object in the background template image is replaced with a sample object to determine the reference image;
[0408] Based on a predetermined random number generation model, a random number subject to a Gaussian distribution is generated;
[0409] For each pixel point in the reference image, the random number is added to the pixel value of the pixel point to obtain the noise reference image.
[0410] Optionally, the image generation model includes a diffusion network and a denoising network;
[0411] The prediction unit 3040 includes:
[0412] A compression module (not shown) is configured to compress the sample spliced image features to obtain sample compressed image features;
[0413] A diffusion module (not shown) is configured to diffuse the sample compressed image features based on the diffusion network to obtain a sample latent space feature vector;
[0414] A denoising module (not shown) is configured to denoise the sample latent space feature vector through the denoising network based on the denoising network control information to obtain a noise prediction result.
[0415] Optionally, the denoising network includes an up-sampling attention network and a down-sampling attention network; and the denoising network control information includes up-sampling network control information and down-sampling network control information.
[0416] The denoising module (not shown) is configured to:
[0417] The down-sampling network control information is fused into a first attention matrix of the down-sampling attention network to update the first attention matrix, and the up-sampling network control information is fused into a second attention matrix of the up-sampling attention network to update the second attention matrix;
[0418] The sample latent space feature vector is denoised through the down-sampling attention network after updating the first attention matrix and the up-sampling attention network after updating the second attention matrix to obtain the noise prediction result.
[0419] Optionally, the sample object image is generated by:
[0420] A sample image including a sample object is obtained;
[0421] perform image segmentation on the sample image based on the preset object segmentation model to obtain a sample segmentation image having a sample object;
[0422] perform image enhancement on the sample segmentation image to obtain a sample object image.
[0423] Optionally, the template mask image is generated in the following manner:
[0424] determining an object contour region of the reference object in the background template image;
[0425] replacing pixel values of each pixel point in the object contour region with a first value and replacing pixel values of each pixel point outside the object contour region with a second value in the background template image to obtain the template mask image.
[0426] Optionally, the contour feature of the reference object in the background template image is determined in the following manner:
[0427] performing object detection on the background template image to obtain an object skeleton map of the reference object;
[0428] extracting pose features from the object skeleton map to obtain a plurality of object pose key points;
[0429] determining the contour feature based on the plurality of object pose key points.
[0430] FIG. 31 is a structural schematic diagram of an image generation apparatus 3100 provided by an embodiment of the present disclosure. The image generation apparatus 3100 comprises:
[0431] a second acquisition unit 3110 configured to acquire an object image comprising a first object, a background image, and description information, wherein the background image contains a second object, and the description information is used to describe replacement from the second object to the first object;
[0432] a third determination unit 3120 configured to determine a splicing image feature based on the background image, a preset noise image, and a mask image, wherein the mask image is obtained by masking the second object in the background image;
[0433] a fourth determination unit 3130 configured to determine denoising control information of an image generation model based on a contour feature of the second object in the background image, the description information, and the object image, wherein the image generation model is generated by the above-mentioned image generation model training method;
[0434] an image generation unit 3140 configured to perform image generation based on the splicing image feature and the denoising control information through the image generation model to obtain a target image, wherein the target image is used to indicate a result of replacing the first object of the object image with the second object in the background image.
[0435] Optionally, the third determining unit 3120 is configured to:
[0436] performing first encoding processing on the preset noise image to obtain noise image features;
[0437] performing second encoding processing on the background image to obtain background image features;
[0438] performing third encoding processing on the mask image to obtain mask image features;
[0439] performing splicing on the noise image features, the background image features, and the mask image features to obtain spliced image features.
[0440] Optionally, the image generation model comprises a diffusion network, a denoising network, and a decoding network.
[0441] The image generation unit 3140 is configured to:
[0442] performing compression processing on the spliced image features to obtain compressed image features;
[0443] performing diffusion processing on the compressed image features based on the diffusion network to obtain an implicit space feature vector;
[0444] performing denoising processing on the implicit space feature vector based on the denoising control information through the denoising network to obtain a denoising result;
[0445] performing feature decoding on the denoising result based on the decoding network to obtain a target image.
[0446] Referring to FIG. 32, FIG. 32 is a structural block diagram of part of a terminal implementing an image generation model training method or an image generation method according to the embodiments of the present disclosure. The terminal can be the object terminal shown in FIG. 1. The terminal includes a radio frequency (RF) circuit 3210, a memory 3215, an input unit 3230, a display unit 3240, a sensor 3250, an audio circuit 3260, a wireless fidelity (WiFi) module 3270, a processor 3280, and a power supply 3290, and the like. Those skilled in the art can understand that the structure of the terminal shown in FIG. 32 does not constitute a limitation to the mobile phone or computer, and can include more or fewer components than those shown, or combine some components, or different component arrangements.
[0447] The RF circuit 3210 can be used for receiving and sending signals in the process of information or call, especially, receiving the downlink information of the base station and processing by the processor 3280; in addition, sending the uplink data to the base station.
[0448] The memory 3215 can be used to store software programs and modules, and the processor 3280 can execute various function applications and data processing of the object terminal by running the software programs and modules stored in the memory 3215.
[0449] The input unit 3230 can be used to receive inputted digital or character information, and to generate key signal input related to the setting and function control of the object terminal. Specifically, the input unit 3230 can include a touch panel 3231 and other input devices 3232.
[0450] The display unit 3240 can be used to display inputted information or provided information and various menus of the object terminal. The display unit 3240 can include a display panel 3241.
[0451] The audio circuit 3260, the speaker 3261, and the microphone 3262 can provide an audio interface.
[0452] In the embodiment, the processor 3280 included in the terminal can execute the training method of the image generation model or the image generation method of the previous embodiments.
[0453] FIG. 33 is a structural block diagram of a part of a server for implementing the training method of the image generation model or the image generation method of the embodiments of the present disclosure. The server can be the image processing server shown in FIG. 1. The server can be greatly different due to different configurations or performances, and can include one or more central processing units (CPUs) 3322 (for example, one or more processors) and a memory 3332, one or more storage media 3330 (for example, one or more mass storage devices) storing application programs 3342 or data 3344. Among them, the memory 3332 and the storage media 3330 can be temporary storage or persistent storage. The programs stored in the storage media 3330 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Further, the central processing unit 3322 can be configured to communicate with the storage media 3330 and execute the series of instruction operations in the storage media 3330 on the server.
[0454] The server can also include one or more power supplies 3326, one or more wired or wireless network interfaces 3350, one or more input / output interfaces 3358, and / or one or more operating systems 3341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0455] The central processor 3322 in the server can be configured to execute the training method of the image generation model or the image generation method according to the embodiments of the present disclosure.
[0456] The embodiments of the present disclosure also provide a computer readable storage medium for storing a computer program, the computer program being configured to execute the training method of the image generation model or the image generation method according to the above embodiments.
[0457] The embodiments of the present disclosure also provide a computer program product comprising a computer program. The processor of the electronic device reads the computer program and executes it, so that the electronic device executes the training method of the image generation model or the image generation method according to the above embodiments.
[0458] The terms "first", "second", "third", "fourth" and the like in the description of the present disclosure and the above drawings, if any, are used to distinguish similar objects, and do not necessarily have to be used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprise" and "include" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0459] It should be understood that in the present disclosure, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases of only A, only B and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b and c can be single or multiple.
[0460] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is two or more, greater than, less than, more than, etc. are not included in the number, above, below, etc. are included in the number.
[0461] In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is merely logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0462] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.
[0463] In addition, each functional unit in the various embodiments of the present disclosure can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of software functional units.
[0464] If the integrated unit is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present disclosure essentially or the part that makes a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present disclosure. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.
[0465] It should also be understood that the various embodiments provided by the present disclosure can be combined in any manner to achieve different technical effects.
[0466] The above is a specific explanation of the embodiments of the present disclosure, but the present disclosure is not limited to the above-described embodiments, and those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present disclosure, and these equivalent modifications or substitutions are included in the scope defined by the claims of the present disclosure.
Claims
1. A training method of an image generation model, executed by an electronic device, the training method comprising: obtaining a plurality of pairs of image-text samples, wherein each pair of image-text samples comprises a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image comprising a sample object, and wherein the noise reference image is obtained by adding noise to a reference image obtained by replacing a reference object in the background template image with the sample object; determining sample spliced image features based on the noise reference image, the background template image, and a template mask image, wherein the template mask image is obtained by masking the reference object in the background template image; determining denoising network control information of the image generation model based on contour features of the reference object in the background template image, the image description information, and the sample object image; performing noise prediction by the image generation model based on the sample spliced image features and the denoising network control information to obtain a noise prediction result corresponding to the noise reference image; and training the image generation model based on comparison results of the noise reference images in the plurality of pairs of image-text samples and the noise prediction results corresponding thereto. The training of the image generation model based on the comparison results of the noise reference images in the plurality of pairs of image-text samples and the noise prediction results corresponding thereto comprises: for each pair of image-text samples, obtaining reference noise in the noise reference image therein, and calculating a sub-loss function of the pair of image-text samples based on a comparison result of the reference noise and the noise prediction result corresponding to the noise reference image, wherein the reference noise is the noise added in the reference image; determining a total loss function based on the sub-loss function of each pair of image-text samples; and training the image generation model based on the total loss function. The noise prediction result is obtained by prediction at a plurality of prediction time steps. The calculation of the sub-loss function of the pair of image-text samples based on the comparison result of the reference noise and the noise prediction result corresponding to the noise reference image comprises: determining prediction noise at a last prediction time step based on the noise prediction result; performing regular term calculation based on the reference noise and the prediction noise to obtain a regular term calculation result; and determining the sub-loss function based on the regular term calculation result. The determination of the sample spliced image features based on the noise reference image, the background template image, and the template mask image comprises: performing first encoding processing on the noise reference image to obtain reference image encoding features; performing second encoding processing on the background template image to obtain background image encoding features; performing third encoding processing on the template mask image to obtain mask image encoding features; and splicing the reference image encoding features, the background image encoding features, and the mask image encoding features to obtain the sample spliced image features. 2. The method of training an image generating model according to claim 1, wherein, 3. The method of training an image generating model according to claim 2, wherein, 4. The method of training an image generation model according to any one of claims 1 to 3, wherein, 5. The method of training an image generation model according to any one of claims 1 to 4, wherein, The denoising network control information of the image generation model is determined based on the contour feature of the reference object in the background template image, the image description information, and the sample object image, and includes: The image description information is encoded to obtain an image description embedding vector; Feature extraction is performed on the sample object image to obtain sample object feature data; The denoising network control information is generated by a preset control network based on the image description embedding vector, the sample object feature data, and the contour feature of the reference object.
6. The method of training an image generation model according to claim 5, wherein, The image description information is encoded to obtain an image description embedding vector, including: The image description information is segmented to obtain a plurality of description words; A target word is determined from the plurality of description words, and a target word embedding feature corresponding to the target word is found based on a preset dictionary; Each of the other description words except the target word is subjected to word embedding processing to obtain a description word embedding feature of each of the other description words; The target word embedding feature and each of the description word embedding features are integrated into the image description embedding vector.
7. The method of training an image generating model according to claim 5 or 6, wherein, The control network includes a first control sub-network and a second control sub-network; the sample object feature data includes a first sample object feature, a second sample object feature, a third sample object feature, and a fourth sample object feature, wherein the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained by performing respective feature extraction on the sample object image; The denoising network control information is generated by a preset control network based on the image description embedding vector, the sample object feature data, and the contour feature of the reference object, including: The image description embedding vector, the first sample object feature, and the contour feature are input into the first control sub-network to generate control information, obtaining first control information; The image description embedding vector and the second sample object feature are input into the second control sub-network to generate control information, obtaining second control information; Based on the third sample object feature and the fourth sample object feature, an up-sampling network control information is determined; Based on the first control information, the second control information, and the first sample object feature and the second sample object feature, a down-sampling network control information is determined; The up-sampling network control information and the down-sampling network control information are integrated into the denoising network control information.
8. The method of training an image generation model according to any one of claims 1 to 7, wherein, The noise reference image is generated by: Determining a reference object in the background template image to be replaced by a sample object reference image; Based on a predetermined random number generation model, a random number subject to a Gaussian distribution is generated; For each pixel point in the reference image, the random number is added to the pixel value of the pixel point to obtain the noise reference image.
9. The method of training an image generation model according to any one of claims 1 to 8, wherein, The image generation model includes a diffusion network and a denoising network; The noise prediction result corresponding to the noise reference image is obtained by performing noise prediction on the image generation model based on the sample splicing image feature and the denoising network control information, and the noise prediction result includes: The sample splicing image feature is compressed to obtain a sample compressed image feature; The sample compressed image feature is diffused based on the diffusion network to obtain a sample latent space feature vector; The sample latent space feature vector is denoised based on the denoising network control information through the denoising network to obtain the noise prediction result.
10. The method of training an image generation model according to claim 9, wherein, The denoising network includes an up-sampling attention network and a down-sampling attention network; and the denoising network control information includes up-sampling network control information and down-sampling network control information. The noise prediction result is obtained by denoising the sample latent space feature vector through the denoising network based on the denoising network control information, and the denoising includes: The down-sampling network control information is fused into a first attention matrix of the down-sampling attention network to update the first attention matrix, and the up-sampling network control information is fused into a second attention matrix of the up-sampling attention network to update the second attention matrix; The sample latent space feature vector is denoised through the down-sampling attention network after updating the first attention matrix and the up-sampling attention network after updating the second attention matrix to obtain the noise prediction result.
11. The method of training an image generation model according to any one of claims 1 to 10, wherein, The sample object image is generated by: obtaining a sample image including the sample object; performing image segmentation on the sample image based on a preset object segmentation model to obtain a sample segmentation image having the sample object; performing image enhancement on the sample segmentation image to obtain the sample object image.
12. The method of training an image generating model according to any one of claims 1 to 11, wherein, The template mask image is generated by: determining an object contour region of the reference object in the background template image; replacing pixel values of each pixel point in the object contour region with a first value and replacing pixel values of each pixel point outside the object contour region with a second value in the background template image to obtain the template mask image.
13. The method of training an image generating model according to any one of claims 1 to 12, wherein, The contour feature of the reference object in the background template image is determined by: performing object detection on the background template image to obtain an object skeleton map of the reference object; performing posture feature extraction on the object skeleton map to obtain a plurality of object posture key points; determining the contour feature based on the plurality of object posture key points.
14. An image generation method performed by an electronic device, the image generation method comprising: obtaining an object image including a first object, a background image, and description information, wherein the background image contains a second object, and the description information is used to describe replacement from the second object to the first object; determining a splicing image feature based on the background image, a preset noise image, and a mask image, wherein the mask image is obtained by masking the second object in the background image; determine denoising control information of an image generation model based on the contour feature of the second object in the background image, the description information, and the object image, the image generation model being generated according to the training method of the image generation model in any one of claims 1 to 13; generate an image based on the spliced image feature and the denoising control information through the image generation model to obtain a target image, wherein the target image is used to indicate a result of replacing the first object in the object image with the second object in the background image.
15. The image generation method of claim 14, wherein, The image generation model comprises a diffusion network, a denoising network, and a decoding network. The image generation model comprises a diffusion network, a denoising network, and a decoding network. The image generation model comprises a diffusion network, a denoising network, and a decoding network. The image generation model comprises a diffusion network, a denoising network, and a decoding network. The image generation model comprises a diffusion network, a denoising network, and a decoding network. The image generation model comprises a diffusion network, a denoising network, and a decoding network.
16. An apparatus for training an image generating model, wherein, The training device of the image generation model comprises: A first acquisition unit is configured to acquire a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image comprising a sample object, and the noise reference image is obtained by adding noise to a reference object in the background template image. A first determination unit is configured to determine a sample spliced image feature based on the noise reference image, the background template image, and a template mask image, wherein the template mask image is obtained by masking the reference object in the background template image. A second determination unit is configured to determine denoising network control information of the image generation model based on a contour feature of the reference object in the background template image, the image description information, and the sample object image. A prediction unit is configured to perform noise prediction on the noise reference image corresponding to the sample spliced image feature and the denoising network control information through the image generation model to obtain a noise prediction result. A training unit is configured to train the image generation model based on a comparison result of the noise reference image in each of the plurality of image-text sample pairs and the noise prediction result corresponding thereto.
17. An image generation apparatus, wherein The image generation device comprises: A second acquisition unit is configured to acquire an object image comprising a first object, a background image, and description information, wherein the background image comprises a second object, and the description information is used to describe replacement from the second object to the first object. A third determination unit is configured to determine a spliced image feature based on the background image, a preset noise image, and a mask image, wherein the mask image is obtained by masking the second object in the background image. a fourth determining unit, configured to determine denoising control information of an image generation model based on the contour feature of the second object in the background image, the description information, and the object image, the image generation model being generated according to the training method of the image generation model in any one of claims 1 to 13; an image generation unit, configured to perform image generation based on the spliced image feature and the denoising control information by using the image generation model, to obtain a target image, wherein the target image is used to indicate a result of replacing the second object in the background image by the first object in the object image.
18. An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to implement the training method of the image generation model according to any one of claims 1 to 13 or the image generation method according to any one of claims 14 to 15.
19. A computer readable storage medium, said storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the training method of the image generation model according to any one of claims 1 to 13 or the image generation method according to any one of claims 14 to 15.
20. A computer program product, comprising a computer program, which is read and executed by a processor of an electronic device, so that the electronic device performs the training method of the image generation model according to any one of claims 1 to 13 or the image generation method according to any one of claims 14 to 15.
Citation Information
Patent Citations
Image generation method and training method and device of image generation model
CN118015144A
Training method and device with text image generation network, electronic equipment and medium
CN118298192A
Figure graph model training method, image prediction method, device, equipment and medium
CN118429755A
Training method of image generation model, related device and medium
CN118570054A
Model Training Method and Apparatus, Text Image Processing Method, Device and Medium
US20250037443A1