Training method for image generation model, and related apparatus and medium

By constructing an image generation model, utilizing background template images, noisy reference images, and sample object images, and combining denoising network control information, the problem of low accuracy of target images in existing technologies is solved, and higher quality image generation is achieved.

WO2026031761A9PCT designated stage Publication Date: 2026-05-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-06-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies for replacing objects in background images are often limited by factors such as the angle between the target object and the background image, lighting filters, etc., resulting in low accuracy of the target image generated by the model and difficulty in collecting enough training data that meets expectations.

Method used

By constructing an image generation model, using background template images, noisy reference images, image description information, and sample object images, the model generates sample stitched image features. It also introduces denoising network control information to perform noise prediction and model training, thereby improving the accuracy of image generation.

Benefits of technology

It improves the accuracy of the image generation model in generating target images, ensures the consistency of lighting and angle of the target object in the background image, and enhances the information content of the training data and the constraints of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025100081_15052026_PF_FP_ABST
    Figure CN2025100081_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a training method for an image generation model, and a related apparatus and a medium. The method comprises: acquiring a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information of the noise reference image, and a sample object image comprising a sample object; on the basis of the noise reference image, the background template image and a template mask image, determining sample splicing image features; on the basis of contour features of a reference object in the background template image, the image description information, and the sample object image, determining denoising network control information; on the basis of the sample splicing image features and the denoising network control information, performing noise prediction by means of an image generation model, so as to obtain a noise prediction result corresponding to the noise reference image; and on the basis of comparison results between the noise reference images in the plurality of image-text sample pairs and the noise prediction results respectively corresponding thereto, training the image generation model. The present disclosure can improve the accuracy of generating a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Training methods, related devices and media for image generation models

[0001] This application claims priority to Chinese Patent Application No. 2024110606399, filed on August 5, 2024, entitled “Training Method, Related Apparatus and Medium for Image Generation Model”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of artificial intelligence technology, and in particular to a training method, related apparatus and medium for an image generation model. Background Technology

[0003] Currently, in various business scenarios such as video production and virtual display, it is often necessary to replace background and object images to create personalized images. For example, when creating an image of item A, it is necessary to replace item B in a background image with item A. Related technologies often use neural network models to extract a first object from a specified background image and create a second object with the same orientation and shooting angle as the first object in the background image. This second object is then embedded into the background image as a texture, allowing it to be displayed within the specified background image.

[0004] However, the effect of the target image generated by the above method is often limited by various factors such as the angle, lighting and filter of the target object (i.e. the second object mentioned above) and the background image. In actual model training, it is often difficult to collect enough training data that meets the expectations, which will affect the model training effect and result in the target image generated by the model not meeting the requirements (for example, the lighting and filter of the target object in the image cannot be consistent with the background image), thus resulting in low accuracy of the target image generated by the model. Summary of the Invention

[0005] This disclosure provides a training method, related apparatus, and medium for an image generation model, which can improve the accuracy of the image generation model in generating target images.

[0006] According to one aspect of this disclosure, a method for training an image generation model is provided, executed by an electronic device, the training method comprising:

[0007] Multiple image-text sample pairs are obtained, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including the sample object. The noise reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object.

[0008] Based on the noise reference image, the background template image, and the template mask image, the features of the sample stitched image are determined, wherein the template mask image is obtained by masking the reference object in the background template image;

[0009] Based on the contour features of the reference object in the background template image, the image description information, and the sample object image, the denoising network control information of the image generation model is determined;

[0010] Based on the features of the sample stitched image and the control information of the denoising network, noise prediction is performed through the image generation model to obtain the noise prediction result corresponding to the noise reference image;

[0011] The image generation model is trained based on the comparison results between the noise reference image and its corresponding noise prediction result in multiple image-text sample pairs.

[0012] According to one aspect of this disclosure, an image generation method is provided, performed by an electronic device, the image generation method comprising:

[0013] Acquire an object image containing a first object, a background image, and descriptive information, wherein the background image contains a second object, and the descriptive information is used to describe the replacement from the second object to the first object;

[0014] Based on the background image, the preset noise image, and the mask image, the features of the stitched image are determined, wherein the mask image is obtained by masking the second object in the background image;

[0015] Based on the contour features of the second object in the background image, the description information, and the object image, the denoising control information of the image generation model is determined, and the image generation model is generated according to the above-described image generation model training method;

[0016] Based on the stitched image features and the denoising control information, an image is generated using the image generation model to obtain a target image, wherein the target image is used to indicate the result of replacing the second object in the background image with the first object in the target image.

[0017] According to one aspect of this disclosure, a training apparatus for an image generation model is provided, the training apparatus for the image generation model comprising:

[0018] The first acquisition unit is used to acquire multiple image-text sample pairs, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including a sample object, wherein the noise reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object;

[0019] The first determining unit is configured to determine the features of the sample stitched image based on the noise reference image, the background template image, and the template mask image, wherein the template mask image is obtained by masking the reference object in the background template image;

[0020] The second determining unit is used to determine the denoising network control information of the image generation model based on the contour features of the reference object in the background template image, the image description information, and the sample object image;

[0021] The prediction unit is used to perform noise prediction based on the features of the sample stitched image and the control information of the denoising network, through the image generation model, to obtain the noise prediction result corresponding to the noise reference image;

[0022] The training unit is used to train the image generation model based on the comparison results between the noise reference image and its corresponding noise prediction result in multiple image-text sample pairs.

[0023] According to one aspect of this disclosure, an image generating apparatus is provided, the image generating apparatus comprising:

[0024] The second acquisition unit is used to acquire an object image including a first object, a background image, and descriptive information, wherein the background image includes a second object, and the descriptive information is used to describe the replacement from the second object to the first object;

[0025] The third determining unit is used to determine the features of the spliced ​​image based on the background image, the preset noise image, and the mask image, wherein the mask image is obtained by masking the second object in the background image;

[0026] The fourth determining unit is used to determine the denoising control information of the image generation model based on the contour features of the second object in the background image, the description information, and the object image, wherein the image generation model is generated by the above-mentioned image generation model training method;

[0027] An image generation unit is configured to generate an image based on the stitched image features and the denoising control information, using the image generation model, to obtain a target image, wherein the target image is used to indicate the result of replacing the second object in the background image with the first object of the object image.

[0028] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the training method or image generation method of the image generation model as described above.

[0029] According to one aspect of this disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the training method or image generation method of the image generation model as described above.

[0030] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that is read and executed by a processor of an electronic device, causing the electronic device to perform the training method or image generation method of the image generation model as described above.

[0031] In this embodiment, when training the image generation model, image-text sample pairs are constructed using a background template image, a noisy reference image, image description information corresponding to the noisy reference image, and a sample object image including the sample object. The noisy reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object. This method allows the reference image used as a reference to better reflect reality. Next, the image features of the background template image, the object mask image corresponding to the background template image, and the noisy reference image are integrated into sample stitched image features. These sample stitched image features are used as input for model training. Since the sample stitched image features possess various image feature information, using them as training data enriches the information content of the training data. Furthermore, this embodiment also introduces image description information (which indicates the background and object in the noisy reference image), contour features of the reference object in the background template image (which reflects the action, posture, etc. of the reference object), and sample object images including the sample object (which reflects the contour, appearance, etc. of the sample object) to jointly generate denoising network control information, enabling the denoising network control information to have multiple constraints. Furthermore, the denoising network control information and the features of the sample stitched image are input into the image generation model. This restricts the noise prediction process based on the denoising network control information, allowing the model to refine its noise prediction of the reference image. Finally, by comparing the differences in noise levels between the predicted noise and the reference noise image, the image generation model is trained to meet the training requirements. This approach improves the accuracy of the generated images.

[0032] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0033] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0034] Figure 1 is a system architecture diagram of the training method of the image generation model and the system application of the image generation method according to an embodiment of the present disclosure.

[0035] Figure 2A is one of the schematic diagrams illustrating the application of the image generation method according to an embodiment of the present disclosure in a video production scenario;

[0036] Figure 2B is a second schematic diagram of the image generation method according to an embodiment of the present disclosure applied in a video production scenario;

[0037] Figure 2C is a third schematic diagram of the image generation method according to an embodiment of the present disclosure applied in a video production scenario;

[0038] Figure 3 is a flowchart of a training method for an image generation model according to an embodiment of the present disclosure;

[0039] Figure 4 is a flowchart of determining a sample object image according to an embodiment of the present disclosure;

[0040] Figure 5 is a flowchart of determining a noise reference image according to an embodiment of the present disclosure;

[0041] Figure 6 is a schematic diagram illustrating the process of determining a noise reference image according to an embodiment of this disclosure;

[0042] Figure 7 is a flowchart of generating sample stitched image features according to an embodiment of the present disclosure;

[0043] Figure 8 is a flowchart of determining a template mask image according to an embodiment of the present disclosure;

[0044] Figure 9A is one of the schematic diagrams illustrating the process of determining the template mask image according to an embodiment of this disclosure;

[0045] Figure 9B is a second schematic diagram illustrating the implementation process of determining the template mask image according to an embodiment of this disclosure;

[0046] Figure 10 is a flowchart of generating denoised network control information according to an embodiment of the present disclosure;

[0047] Figure 11 is a flowchart of determining the contour features of a reference object in a background template image according to an embodiment of the present disclosure;

[0048] Figure 12 is a schematic diagram of the implementation process of key point extraction when determining contour features according to an embodiment of the present disclosure;

[0049] Figure 13 is a flowchart of generating an image description embedding vector according to an embodiment of the present disclosure;

[0050] Figure 14 is a schematic diagram of the process of generating an image description embedding vector according to an embodiment of the present disclosure;

[0051] Figure 15 is a flowchart of generating denoised network control information according to an embodiment of the present disclosure;

[0052] Figure 16 is a schematic diagram of the implementation process of generating denoised network control information according to an embodiment of the present disclosure;

[0053] Figure 17 is a flowchart of generating noise prediction results according to an embodiment of the present disclosure;

[0054] Figure 18 is a schematic diagram of the process of generating noise prediction results according to an embodiment of the present disclosure;

[0055] Figure 19 is a flowchart of a noise reduction process according to an embodiment of the present disclosure;

[0056] Figure 20 is a schematic diagram illustrating the implementation process of the downsampling process during denoising processing according to an embodiment of the present disclosure;

[0057] Figure 21 is a schematic diagram illustrating the implementation process of the upsampling process during denoising according to an embodiment of the present disclosure;

[0058] Figure 22 is a flowchart of a training image generation model according to an embodiment of the present disclosure;

[0059] Figure 23 is a flowchart of determining a sub-loss function according to an embodiment of the present disclosure;

[0060] Figure 24 is a schematic diagram illustrating the implementation details of a training image generation model according to an embodiment of the present disclosure;

[0061] Figure 25 is a flowchart of an image generation method according to an embodiment of the present disclosure;

[0062] Figure 26 is a flowchart of determining target stitched image features according to an embodiment of the present disclosure;

[0063] Figure 27 is a flowchart of generating a target image according to an embodiment of the present disclosure;

[0064] Figure 28 is a simplified schematic diagram of an embodiment of an image generation method according to the present disclosure;

[0065] Figure 29 is a schematic diagram illustrating the implementation details of an image generation method according to an embodiment of the present disclosure;

[0066] Figure 30 is a block diagram of a training apparatus for an image generation model according to an embodiment of the present disclosure;

[0067] Figure 31 is a block diagram of an image generation apparatus according to an embodiment of the present disclosure;

[0068] Figure 32 is a terminal structure diagram of a training method for an image generation model according to an embodiment of the present disclosure;

[0069] Figure 33 is a server structure diagram of a training method for an image generation model according to an embodiment of the present disclosure. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0071] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:

[0072] The Cross Attention Control Module is a module used in deep learning to establish a cross-attention mechanism among multiple inputs. Specifically, the Cross Attention Control Module helps the model automatically learn the correlations between inputs when processing multiple inputs, thereby improving model performance.

[0073] The system architecture and scenarios in which this disclosure is applied are described below.

[0074] Figure 1 is a system architecture diagram of a training method for an image generation model according to an embodiment of the present disclosure, and an image generation method applied thereto. It includes an object terminal 140, an Internet 130, a gateway 120, an image processing server 110, and an image database 150, etc.

[0075] The target terminal 140 includes various forms such as desktop computers, laptops, PDAs (personal digital assistants), tablets, mobile phones, in-vehicle terminals, home theater terminals, smart TVs, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. The target terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data. The target terminal 140 includes an image processing system, which receives a selected background image and a target image, and submits them to an image processing server 110 so that the image processing server 110 can generate a target image with the target object in the target image as the foreground and the background image as the background.

[0076] Image processing server 110 refers to a computer system that can provide certain services to object terminal 140. Compared with ordinary object terminal 140, image processing server 110 has higher requirements in terms of stability, security, and performance. Image processing server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a high-performance computer (e.g., a virtual machine), a combination of portions of multiple high-performance computers (e.g., virtual machines), or a cloud server, etc. Image processing server 110 includes various types of services, and the implementation of each service of image processing server 110 is often associated with some intermediate databases or storage media. Image processing server 110 is used to generate target images with the target object in the object image as the foreground and the background image as the background using a trained image generation model. Image database 150 is used to store various images such as target images, object images, and background images.

[0077] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from object terminal 140 to image processing server 110 are forwarded to the corresponding image processing server 110 via gateway 120. Similarly, messages sent from image processing server 110 to object terminal 140 are also forwarded to the corresponding object terminal 140 via gateway 120.

[0078] The embodiments disclosed herein can be applied to various scenarios, such as the video production scenarios shown in Figures 2A-2C.

[0079] As shown in Figure 2A, when an object needs to create an image with object A as the foreground and background image 5 as the background during video production, the object will log into the image processing system on its terminal and enter the image processing workflow. At this time, a prompt field will appear on the page: "Please select a background image, upload the target object image, and enter an image description," providing editing areas for selecting the background image, selecting the target object image, and entering an image description. Based on this, the object selects "Background Image 5" in the background image selection area; selects "D:\Object A's Object Image\Image 1 + Image 2" in the target object image selection area; and enters "Replace Object B in Background Image 5 with Object A" in the image description input area, and clicks the "OK" button to confirm that object B in background image 5 will be replaced with object A using background image 5, image 1 of object A, and image 2 of object A.

[0080] As shown in Figure 2B, after clicking the "OK" button on the object, a prompt window will be displayed on the page. The prompt window contains the prompt field "Generating the mask image corresponding to background image 5 and the object outline image of object B, and calling the image generation model to perform image synthesis based on the mask image and object outline image, etc. Please wait patiently...", to illustrate the image synthesis process of the image generation model to the object.

[0081] As shown in Figure 2C, once the image generation model has finished executing, the page will display the message "Object B in background image 5 has been replaced with object A, see the image below for details. Furthermore, the generated image has been saved to the path: D:\Composite Image Folder\Video Production.", and will show a comparison between background image 5 and the generated target image. Background image 5 contains bottle A (object B), and the target image contains bottle B (object A). Through the image generation model in the image processing system, bottle A in background image 5 was replaced with bottle B while keeping the background unchanged.

[0082] It should be noted that the image generation method of this disclosure can be applied not only to the video production scenario described above, but also to various other application scenarios such as product display in e-commerce and personalized image creation in social media.

[0083] The training method of the image generation model according to the embodiments of this disclosure will be described in general below.

[0084] According to one embodiment of this disclosure, a method for training an image generation model is provided.

[0085] The training method for this image generation model is generally applied in business scenarios where reference objects (people, animals, items, etc.) in a fixed background image need to be replaced with target objects (people, animals, items, etc.), such as the video production scenario and item display scenario shown in Figures 2A-2C. This disclosure provides a scheme for model training based on image description and the differences between the objects to be generated and the background image, which can improve the accuracy of the image generation model in generating target images.

[0086] As shown in Figure 3, the training method of the image generation model according to an embodiment of the present disclosure can be executed by an electronic device, which may be the image processing server or object terminal shown in Figure 1. The training method of the image generation model may include:

[0087] Step 310: Obtain multiple image-text sample pairs;

[0088] Step 320: Determine the features of the sample stitched image based on the noisy reference image, background template image, and template mask image;

[0089] Step 330: Based on the contour features of the reference object in the background template image, the image description information, and the sample object image, determine the denoising network control information of the image generation model;

[0090] Step 340: Based on the features of the sample stitched image and the control information of the denoising network, noise prediction is performed through the image generation model to obtain the noise prediction result corresponding to the noise reference image;

[0091] Step 350: Train the image generation model based on the comparison results of the noisy baseline image and its corresponding noise prediction results in multiple image-text sample pairs.

[0092] Steps 310-350 are described in detail below.

[0093] In step 310, multiple image-text sample pairs are obtained.

[0094] In this embodiment of the disclosure, an image-text sample pair is used as training data, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including the sample object.

[0095] Background template image refers to an image that provides background content for the final image to be generated. In this embodiment of the disclosure, sample objects need to be added to the background of the background template image. The background template image also includes a reference object that will be replaced by the sample object in the sample object image.

[0096] A sample object image that includes a sample object refers to an image that reflects the characteristics of the sample object (object posture, object action, object appearance, etc.). Each image-text sample pair may include one or more sample object images.

[0097] The noisy reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object. In other words, the reference object in the background template image can be replaced with the sample object to obtain the reference image, and then noise can be iteratively added to the reference image to obtain the noisy reference image.

[0098] The image description information corresponding to the noisy reference image is used to indicate the constraints on the denoising process. For example, the image description information could be that person A is generated in the background template image.

[0099] In step 320, the features of the sample stitched image are determined based on the noise reference image, the background template image, and the template mask image.

[0100] A template mask image is obtained by masking a reference object in a background template image.

[0101] The sample stitched image features are used to indicate the stitching result of the features of the noise reference image, background template image, and template mask image, respectively.

[0102] In this specific implementation, firstly, the noise reference image, background template image, and template mask image are mapped from pixel space to latent vector space to obtain vector features corresponding to the noise reference image, the background template image, and the template mask image, respectively. Next, the vector features corresponding to the noise reference image, the background template image, and the template mask image are concatenated to obtain the sample concatenated image features.

[0103] In step 330, the denoising network control information of the image generation model is determined based on the contour features of the reference object in the background template image, the image description information, and the sample object image.

[0104] The contour features of the reference object in the background template image are used to indicate the contour characteristics of the reference object's actions, postures, etc. in the background template image.

[0105] The image generation model of this disclosure is a neural network model built based on the stable diffusion (SD) model. This image generation model typically takes a background image, an object image, and information describing the object in the object image that replaces the background image as input, and outputs a composite image with the scene in the background image as the background and the object in the object image as the foreground. The image generation model can embed objects into specific locations in an image, meeting image generation needs in various scenarios.

[0106] The control information of the denoising network is used as a conditional constraint to assist the image generation model in image denoising.

[0107] In the specific implementation of this embodiment, the image description information can be encoded to obtain the corresponding image description feature vector, and the sample object image can be processed by feature extraction to obtain the corresponding sample object feature vector. Then, the denoising network control information can be determined based on the contour features of the reference object in the background template image, as well as the above-mentioned image description feature vector and sample object feature vector.

[0108] To save space, the specific process of determining the denoising network control information based on the contour features, image description information, and sample object image of the reference object in the background template image of this disclosure will be described in more detail below, and will not be repeated here.

[0109] In step 340, based on the features of the sample stitched image and the control information of the denoising network, noise prediction is performed through the image generation model to obtain the noise prediction result corresponding to the noise reference image.

[0110] The noise prediction results are used to indicate the amount of noise contained in the image features generated by the image generation model that are consistent with the noise reference image.

[0111] In the specific implementation of this embodiment, the features of the sample stitched image can be used as the object to be denoised by the image generation model, and the control information of the denoising network can be used as the control information referenced by the image generation model when performing denoising processing. When performing denoising processing, the image generation model can iteratively perform denoising processing on the features of the sample stitched image based on the control information of the denoising network, so as to obtain the noise prediction result of the noise reference image in each round of denoising processing.

[0112] To save space, the specific process of noise prediction based on sample stitched image features and denoising network control information through an image generation model in this embodiment will be described in more detail below, and will not be repeated here.

[0113] In step 350, an image generation model is trained based on the comparison results between the noisy reference image and its corresponding noise prediction result in multiple image-text sample pairs.

[0114] In this specific implementation, the model parameters of the image generation model are adjusted according to the degree of noise difference between the noise contained in the noise reference image of the image-text sample pair and the noise prediction result. Steps 310-350 above are repeated to continuously reduce the noise difference between the noise contained in the noise reference image of the image-text sample pair and the noise prediction result until the noise difference between the noise contained in the noise reference image of multiple image-text sample pairs and the noise prediction result meets the model training requirements. At this point, the update of the model parameters of the image generation model is stopped, and the model parameters at this time are taken as the final model parameters. The image generation model with the final model parameters is taken as the trained image generation model.

[0115] Through steps 310-350 above, in this embodiment of the disclosure, when training the image generation model, image-text sample pairs are constructed using a background template image, a noisy reference image, image description information corresponding to the noisy reference image, and a sample object image of the sample object. The noisy reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object. This method allows the reference image used as a reference to be more closely aligned with reality. Next, the image features of the background template image, the template mask image of the background template image (obtained by masking the reference image in the background template image), and the noisy reference image are integrated into sample stitched image features. These sample stitched image features are used as input for model training. Since the sample stitched image features possess a variety of image feature information, using them as training data can enrich the information content of the training data. Furthermore, this embodiment introduces image description information (which indicates the background and objects in the noisy reference image), contour features of the reference object in the background template image (which reflects the action, posture, etc. of the reference object), and sample object images of the sample object (which reflect the contour, appearance, etc. of the sample object) to jointly generate denoising network control information, giving the denoising network control information multiple constraints. Further, the denoising network control information and sample stitched image features are input into the image generation model, so that when the image generation model performs noise prediction on the sample stitched image features, it is constrained by the denoising network control information, and the noise prediction process is corrected according to the denoising network control information, making the noise prediction result of the noisy reference image output by the image generation model more accurate. Finally, the image generation model is trained by comparing the noise difference between the noise prediction result and the noise contained in the noisy reference image, thereby obtaining an image generation model that meets the training requirements. This method can improve the accuracy of the model in generating the target image.

[0116] The above is a general description of steps 310-350. The specific implementations of steps 310, 320, 330, 340 and 350 will be described in detail below.

[0117] Step 310 will be described in detail below.

[0118] In step 310, multiple image-text sample pairs are obtained. Each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including the sample object. The noise reference image is obtained by adding noise to the reference image obtained by replacing the reference object in the background template image with the sample object.

[0119] Referring to Figure 4, in one embodiment, the sample object image of the sample object is determined in the following manner:

[0120] Step 410: Obtain a sample image containing the sample object;

[0121] Step 420: Based on the preset object segmentation model, perform image segmentation on the sample image to obtain a sample segmentation image with sample objects;

[0122] Step 430: Perform image enhancement on the sample segmentation image to obtain the sample object image.

[0123] Steps 410-430 are described in detail below.

[0124] In step 410, the sample image is an image containing the sample object.

[0125] In a specific implementation of this embodiment, with authorization, multiple sample images containing sample objects can be extracted from an existing image database, and video data containing sample objects can be extracted from an existing video database. The video data is then segmented into video frames, and each video frame containing the sample object is used as a sample image.

[0126] In step 420, the sample segmentation image refers to a local image containing the sample object segmented from the sample image; the sample segmentation image is a part of the sample image.

[0127] The object segmentation model can be a lightweight semantic segmentation model such as BiSeNet-v2 or PP_LiteSeg.

[0128] Taking the semantic segmentation model PP_LiteSeg as an example, the object segmentation model includes an encoding module, a pyramid pooling module, a decoding module, and an attention fusion module. Specifically, firstly, the sample image is input into the encoding module of the object segmentation model. The encoding module performs multi-scale encoding on the sample image, sequentially obtaining a first sample image feature (one-quarter the size of the sample image), a second sample image feature (one-eighth the size of the sample image), a third sample image feature (one-sixteenth the size of the sample image), and a fourth sample image feature (one-thirty-second the size of the sample image). Next, the pyramid pooling module performs feature pooling on the fourth sample image feature to obtain the pooled sample image feature. Further, the attention fusion module fuses the pooled sample image feature and the third sample image feature to obtain the first image fusion feature. Then, the attention fusion module fuses the first image fusion feature and the second sample image feature to obtain the second image fusion feature. Finally, the decoding module decodes the second image fusion feature to obtain the sample segmented image containing the sample object.

[0129] In step 430, image enhancement of the sample segmentation image includes, but is not limited to, contrast enhancement, brightness enhancement, sharpening, and noise reduction. Specifically, firstly, when enhancing the sample segmentation image, linear stretching or logarithmic transformation methods are used to expand the pixel value distribution of the sample segmentation image to make the bright and dark areas of the sample segmentation image more distinct, thereby achieving contrast enhancement. Next, after contrast enhancement, the brightness values ​​of all pixels in the sample segmentation image are adjusted to improve the overall brightness of the sample segmentation image. Further, after brightness adjustment, Gaussian filtering or median filtering is used to smooth the sample segmentation image and reduce noise. Finally, after noise reduction, edge enhancement is performed on the sample segmentation image using the Laplacian operator to achieve sharpening processing, resulting in the sample object image.

[0130] The advantage of this embodiment is that by performing image segmentation on sample images containing sample objects, and segmenting local images containing sample objects from the sample images, interference from irrelevant image information can be effectively eliminated, thereby improving the image quality of the sample object images used for training. Furthermore, various image enhancement processes, such as contrast enhancement, brightness enhancement, sharpening, and noise reduction, are applied to the segmented sample images, which can further improve the image quality of the sample object images.

[0131] Referring to Figure 5, in one embodiment, the noise reference image is determined in the following manner:

[0132] Step 510: Determine to replace the reference object in the background template image with the baseline image of the sample object;

[0133] Step 520: Generate random numbers that follow a Gaussian distribution based on a predetermined random number generation model;

[0134] Step 530: For each pixel in the reference image, add a random number to the pixel value of the pixel to obtain the noisy reference image.

[0135] Steps 510-530 are described in detail below.

[0136] In step 510, the reference image is the image data obtained by replacing the reference image in the background template image with the sample object.

[0137] In this specific implementation, image editing software can be used to replace the reference object in the background template image with the sample object, thereby using the resulting replaced image as the base image for replacing the reference object in the background template image with the sample object. The image editing software can be software such as Adobe Photoshop.

[0138] In step 520, the predetermined random number generation model refers to a random number generator, and the random number is a randomly generated number. In this embodiment of the disclosure, the random numbers that follow a Gaussian distribution can be repeated.

[0139] In this specific implementation, the predetermined random number generation model includes a library function (numpy.random.normal()). Specifically, using the library function of the predetermined random number generation model, a random number is generated based on the Gaussian distribution's mean of 0 and standard deviation of 1. This random number follows a Gaussian distribution. The Gaussian-distributed random number can be represented as a random noise graph, the image size of which is the same as the reference image.

[0140] In step 530, for each pixel in the reference image, firstly, the pixel value of the pixel in the reference image is determined, and then a random number corresponding to the pixel in the random noise image is determined based on the pixel's position in the reference image. This random number is used as the noise value corresponding to the pixel. Next, the pixel value and its corresponding noise value are added together to obtain the noisy pixel value of the pixel. Finally, based on the noisy pixel values ​​of each pixel, a noisy reference image is obtained.

[0141] Figure 6 illustrates the specific implementation process of superimposing random noise onto a reference image. The reference image is an 8×8 image with 64 pixels. The random noise map corresponding to random numbers following a Gaussian distribution is also an 8×8 image with 64 pixels. Specifically, for each pixel in the reference image, the pixel value of each pixel is first determined, where the pixel value includes 1, 2, 3, 4, 5, or 6. Next, for each pixel, the corresponding noise value in the random noise map is determined, where the noise value includes 0, 1, 2, 3, 4, 5, or 6. Further, for each pixel, the pixel value and its corresponding noise value are added together to obtain the noise-added pixel value for each pixel. For example, in the first row of the reference image, the first pixel has a pixel value of 1 and a noise value of 5, so the noise-added pixel value is 1 + 5 = 6. The second pixel has a pixel value of 1 and a noise value of 1, so the noise-added pixel value is 1 + 1 = 2. If the pixel value of the third pixel is 1 and the noise value is 4, then the noise-added pixel value is 1+4=5; and so on, until the noise-added pixel value of the last pixel in the last row is determined, thus generating a noise baseline image.

[0142] The advantage of this embodiment is that by adding random noise conforming to a Gaussian distribution to the baseline image obtained by replacing the reference object in the background template image with the sample object, environmental noise is introduced into the idealized baseline image, making the final noisy baseline image more in line with the real situation, thereby improving the realism and accuracy of the noisy baseline image and making the noisy baseline image more referential.

[0143] Step 320 will be described in detail below.

[0144] In step 320, the features of the sample stitched image are determined based on the noise reference image, the background template image, and the template mask image, wherein the template mask image is obtained by masking the reference object in the background template image.

[0145] Referring to Figure 7, in one embodiment, step 320 specifically includes, but is not limited to, the following steps 710-740:

[0146] Step 710: Perform a first encoding process on the noisy reference image to obtain the reference image encoding features;

[0147] Step 720: Perform a second encoding process on the background template image to obtain the background image encoding features;

[0148] Step 730: Perform a third encoding process on the template mask image to obtain the mask image encoding features;

[0149] Step 740: Concatenate the base image coding features, background image coding features, and mask image coding features to obtain the sample concatenated image features.

[0150] Steps 710-740 are described in detail below.

[0151] In step 710, the reference image coding features are used to indicate the result of converting the noisy reference image from the data space to the pixel space.

[0152] In the specific implementation of this embodiment, a preset image encoder can be used to perform a first encoding process on the noisy reference image, converting the noisy reference image from pixel space to latent vector space to capture the key image information of the noisy reference image and obtain the reference image encoding features. The key image information of the noisy reference image includes, but is not limited to, the texture information, edge information, and corner information of the noisy reference image.

[0153] In step 720, the background image encoding features are used to indicate the result of converting the background template image from the data space to the pixel space.

[0154] In this specific implementation, step 720 is similar to step 710 described above. The difference lies in the images to be encoded and the encoder parameters of the image encoders used. For the sake of brevity, these details will not be elaborated further.

[0155] In step 730, the mask image encoding features are used to indicate the result of converting the template mask image from the data space to the pixel space.

[0156] In this specific implementation, step 730 is similar to step 710 described above. The difference lies in the fact that the images to be encoded are different, and the encoder parameters of the image encoders used are different. For the sake of brevity, these details will not be elaborated further.

[0157] In step 740, the reference image coding features, background image coding features, and mask image coding features are concatenated to form a vector feature with a larger number of feature channels. The concatenated vector feature is then determined as the sample concatenated image feature.

[0158] For example, the feature dimensions of the base image encoding features, background image encoding features, and mask image encoding features are all W*H*4, where W is the feature length of the base image encoding features, background image encoding features, and mask image encoding features, H is the feature height of the base image encoding features, background image encoding features, and mask image encoding features, and 4 is the number of feature channels of the base image encoding features, background image encoding features, and mask image encoding features. After the above feature concatenation operation, the feature dimension of the resulting sample concatenated image features is W*H*12.

[0159] The advantage of this embodiment is that it transforms the image information of the noisy reference image, background template image, and template mask image from pixel space to latent vector space. Then, it concatenates the image information (reference image encoding features, background image encoding features, and mask image encoding features) of the noisy reference image, background template image, and template mask image in the latent vector space to form a sample concatenated image feature containing noisy reference image information, template image information, and reference object mask information. Furthermore, using this sample concatenated image feature as input to the image generation model during training allows the model's input data to incorporate image generation effect information, background template information, and reference object information from the background template. This significantly improves the richness and comprehensiveness of the feature information in the sample concatenated image feature, facilitating the training model's learning and mining of various image information, thereby improving the model's image generation accuracy.

[0160] Referring to Figure 8, in one embodiment, the template mask image is determined in the following manner:

[0161] Step 810: Determine the object outline region of the reference object in the background template image;

[0162] Step 820: In the background template image, replace the pixel values ​​of each pixel within the object outline area with the first value, and replace the pixel values ​​of each pixel outside the object outline area with the second value to obtain the template mask image.

[0163] Steps 810-820 are described in detail below.

[0164] In step 810, the object outline region is used to indicate the minimum bounding rectangle region corresponding to the object outline of the reference object.

[0165] In this specific implementation, firstly, each pixel representing the reference object is determined in the background template image. Next, a two-dimensional coordinate system is constructed, with the top-left pixel in the background template image as the origin, the height of the background template image as the vertical axis, the width of the background template image as the horizontal axis, and the distance between two adjacent pixels as one unit length. Further, based on the constructed two-dimensional coordinate system, the coordinate data for each pixel representing the reference object is determined. The maximum horizontal coordinate value x is then selected from the coordinate data of each pixel. max Minimum x-coordinate min , the maximum value of the ordinate y max Minimum value of the ordinate y min Finally, based on the maximum and minimum x-coordinates, the maximum and minimum y-coordinates, the minimum bounding rectangle region corresponding to the object outline of the reference object is determined, and this minimum bounding rectangle region is defined as the object outline region of the reference object. The coordinates of the four endpoints of the minimum bounding rectangle region are (x, y, ..., y). max y max ), (x max y min ), (x min y max ), (x min y min ).

[0166] In step 820, the first value refers to the value 1 that makes the pixel appear white, and the second value refers to the value 0 that makes the pixel appear black.

[0167] In this specific implementation, firstly, the background template image is divided into an object outline region and other regions outside the object outline region. Next, for each pixel within the object outline region, the pixel value is replaced with a first value, making the object outline region appear white. For each pixel outside the object outline region, the pixel value is replaced with a second value, making the object outline region appear black. Thus, based on the changes in the pixel values ​​of each pixel in the background template image, a template mask image is obtained. The template mask image is often represented as a black and white image.

[0168] Figure 9A illustrates a simplified diagram of pixel-level masking. Specifically, in an 8×8 background template image, the minimum bounding rectangle corresponding to the object outline of the reference object is a 4×4 image region. This object outline region contains 3 pixels with a value of 9, 2 pixels with a value of 8, 2 pixels with a value of 1, 2 pixels with a value of 4, and 7 pixels with a value of 2. Based on this, all 16 pixels are replaced with 1, and all pixels in the background template image except these 16 pixels are replaced with 0, resulting in a template mask image where each pixel has a value of either 0 or 1. In this template mask image, the object outline region of the reference object is displayed as white, while all other parts are displayed as black.

[0169] Figure 9B illustrates a simplified diagram of image-level masking. Specifically, the background template image contains a reference object (a person) and several other objects (a pentagram and polygons). Based on this, the minimum bounding rectangle corresponding to the outline of the reference object (the person) is first determined. The pixel values ​​of each pixel within the minimum bounding rectangle are set to 1, while the pixel values ​​of each pixel outside the minimum bounding rectangle in the background template image are set to 0. This process masks the background template image, resulting in a black and white template mask image, where the object outline area is represented by a white rectangular area.

[0170] The advantage of this embodiment is that by determining the object contour region of the reference object in the background template image, the image position of the reference object in the background template image can be clearly determined. By converting the pixel values ​​of each pixel, the background template image is binarized (masked), which can effectively increase the distinguishability of the pixel positions occupied by the reference object and non-reference objects in the background template image. This allows the model to better explore the object contour features of the reference object based on the template mask image, thereby improving the accuracy of replacing the reference object with the target object.

[0171] Step 330 will be described in detail below.

[0172] In step 330, the denoising network control information of the image generation model is determined based on the contour features of the reference object in the background template image, the image description information, and the sample object image.

[0173] Referring to Figure 10, in one embodiment, step 330 specifically includes, but is not limited to, the following steps 1010-1030:

[0174] Step 1010: Encode the image description information to obtain the image description embedding vector;

[0175] Step 1020: Extract features from the sample object image to obtain sample object feature data;

[0176] Step 1030: Based on the image description embedding vector, sample object feature data, and the contour features of the reference object, control information is generated through a preset control network to obtain denoising network control information.

[0177] Steps 1010-1030 are described in detail below.

[0178] In step 1010, the image description embedding vector is used to indicate the vector representation of the conditional constraints on the denoising process in the image description information.

[0179] In this specific implementation, a text encoder can be used to encode the image description information, transforming it from a data space to a vector space to obtain the image description embedding vector. The text encoder includes a word segmenter, an embedding layer, and a text attention calculation module.

[0180] Specifically, first, the image description information is input into a word segmenter, which segments the image description information into multiple descriptive words. Next, an embedding layer performs word embedding processing on each descriptive word, converting each descriptive word into a vector form, resulting in a descriptive word vector for each descriptive word. Finally, a text attention calculation module performs attention calculation on the descriptive word vectors of each descriptive word to obtain the image description embedding vector.

[0181] In step 1020, the sample object feature data is used to indicate the object features of the sample object; the sample object feature data can guide the denoising network of the image generation model to restore the object in the image to be more closely related to the sample object during the denoising process.

[0182] For example, when the sample object is a person, the sample object feature data includes, but is not limited to, indicators of facial contours, facial features, etc.

[0183] It should be noted that the feature data of the sample objects can be obtained through a fine-tuned network generated by a deep learning model-based fine-tuning technique (LoRA technique). The fine-tuned network is a neural network based on the cross-attention algorithm.

[0184] In this specific implementation, firstly, the sample object image is input into the fine-tuning network, which performs linear projection on the sample object image to generate the key vector, value vector, and query vector corresponding to the sample object image. Next, cross-attention calculation is performed using the key vector, value vector, and query vector to obtain the calculation result, which is then converted into vector form to obtain the sample object feature data.

[0185] In step 1030, a preset control network is used to generate control information for the denoising network of the image generation model based on various input data. The input data allowed by the preset control network includes conditional constraints on the image to be generated, the outline of the object in the background template image, the object features of the object to be replaced, etc.

[0186] It should be noted that, in order to improve the accuracy of the denoising network control information output by the control network, the control network in this embodiment of the disclosure can also be a neural network based on the cross-attention algorithm.

[0187] The denoising network control information is used to guide and control the image generation model's attention to and learning of various feature information during the denoising process. The denoising network control information is also used to indicate the degree of fine-tuning of various image feature information during the denoising process.

[0188] To save space, the specific process of generating control information through a preset control network to obtain denoising network control information in the embodiments of this disclosure will be described in detail below, and will not be repeated here.

[0189] The advantage of this embodiment is that it utilizes image description information, the contour features of the reference object in the background template image, and the object features of the sample object to jointly generate the denoising network control information for the image generation model. This allows for denoising based on multiple constraints during the image denoising process, achieving precise fine-tuning and flexible control of image denoising. This approach enables the model to perform image denoising under multi-dimensional constraints, improving the model's image denoising capability. It also helps the model generate image content that better matches the image description and makes the objects in the final generated image more closely resemble the actions and poses of the reference object in the background template image.

[0190] Referring to Figure 11, in one embodiment, the contour features of the reference object in the background template image are determined in the following manner:

[0191] Step 1110: Perform object detection on the background template image to obtain the object skeleton map of the reference object;

[0192] Step 1120: Extract pose features from the object skeleton diagram to obtain multiple object pose key points;

[0193] Step 1130: Determine contour features based on multiple object pose key points.

[0194] Steps 1110-1130 are described in detail below.

[0195] In step 1110, the object skeleton diagram is used to indicate the skeleton outline of the reference object in the background template image.

[0196] In this specific implementation, firstly, a preset detection algorithm is used to locate objects in the background template image to identify the position of the reference object, obtaining object detection results. These object detection results are formed by multiple pixels constituting the reference object. The preset detection algorithm includes, but is not limited to, YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) algorithms. Next, based on the pixels contained in the object detection results, the object skeleton map of the reference object is determined.

[0197] In step 1120, object pose keypoints are used to indicate the action and pose of the reference object in the background template image.

[0198] In this specific implementation, firstly, the object skeleton diagram is input into a preset pose estimation model. Using the preset pose estimation model, the coordinates of pixels representing important parts of the reference object in the object skeleton diagram are located, resulting in a set of key point coordinates. Next, the coordinates of each key point output by the pose estimation model are determined as the object's pose key points. The preset pose estimation model includes, but is not limited to, neural network models based on deep learning algorithms such as PoseNet and AlphaPose.

[0199] In step 1130, firstly, based on the relative positions and relationships of the object's pose keypoints, multiple object pose keypoints are connected to form a complete object line drawing. Next, the object line drawing is converted into an image representation acceptable to the control network to obtain the contour features of the reference object.

[0200] Figure 12 illustrates the keypoint detection process for a person (reference object) in a background template image. Specifically, a keypoint extraction algorithm is used to extract keypoints from the background template image, obtaining facial keypoints and skeletal keypoints of the reference object. Facial keypoints reflect the facial contour and features of the reference object, while skeletal keypoints reflect the pose of the reference object within the background template image. Based on this, the facial keypoints and skeletal keypoints are collectively identified as the object pose keypoints of the reference object.

[0201] The advantage of this embodiment is that, when determining the contour features of the reference object, the reference object is first located in the background template image, and its skeleton features are constructed based on its location. Furthermore, using a pose estimation model to detect key points in the skeleton image improves the efficiency and accuracy of determining pose key points. This allows for the drawing of the reference object's contour features based on the relative positions of multiple pose key points, improving the accuracy of contour feature determination and also enhancing the accuracy of the control information from the denoising network.

[0202] Because conventional encoding methods for converting image description information into embedding vectors often result in inaccuracies in the conditional constraints generated for the image. This disclosure provides a scheme for encoding image description information based on fine-tuning techniques, which improves the accuracy of the generated image description embedding vectors, thereby enhancing the effectiveness of the conditional constraints used to generate control information for denoising networks.

[0203] It should be noted that the encoding of image description information in this embodiment is based on a preset text encoder, which includes a tokenizer, an embedding layer, and a text attention calculation module.

[0204] Referring to Figure 13, in one embodiment, step 1010 specifically includes, but is not limited to, the following steps 1310-1340:

[0205] Step 1310: Segment the image description information to obtain multiple descriptive words;

[0206] Step 1320: Identify the target word among multiple descriptive words, and find the target word embedding feature corresponding to the target word based on the preset dictionary;

[0207] Step 1330: For each of the other descriptive words besides the target word, perform word embedding processing on each other descriptive word to obtain the descriptive word embedding features of each other descriptive word.

[0208] Step 1340: Integrate the target word embedding features and the embedding features of each descriptive word into an image description embedding vector.

[0209] Steps 1310-1340 are described in detail below.

[0210] In step 1310, descriptive words are used to indicate the vocabulary information contained in the image description information.

[0211] In this specific implementation, firstly, the image description information can be input into the word segmenter of the text encoder. Next, the word segmenter will segment the image description information based on spaces, punctuation marks, or delimiters to obtain multiple words. Further, stop words are removed from the segmented words to obtain multiple descriptive words.

[0212] In step 1320, the target word refers to the word in the image description information that needs to be found by word list search to find the corresponding embedded vector representation. The target word embedding feature is used to indicate the vector representation of the target word in the image description information.

[0213] The preset dictionary is set in advance in the embedding layer. The preset dictionary in the embedding layer is used to indicate the correspondence between the index of each descriptor and the word embedding feature.

[0214] In this specific implementation, firstly, for multiple descriptive words, a pre-defined descriptive word to be represented by fixed characters is selected as the target word. Next, the index corresponding to the target word is input into the embedding layer of the text encoder. Based on the index, feature lookup is performed in a pre-defined dictionary within the embedding layer to find the word embedding feature corresponding to the index, which is then used as the target word embedding feature.

[0215] In step 1330, the descriptor embedding features are used to indicate the vector representation of other descriptors in the image description information.

[0216] In this specific implementation, for descriptive words other than the target word among multiple descriptive words, firstly, the index of each other descriptive word is determined based on the segmentation result of the word segmenter. Next, the embedding layer of the text encoder is used to perform word embedding processing on each other descriptive word. The word embedding features corresponding to the index of each other descriptive word are found in the preset dictionary of the embedding layer, thus obtaining the descriptive word embedding features of the other descriptive words.

[0217] In step 1340, firstly, the target word embedding features and descriptive word embedding features are input into the text attention calculation module, which outputs the descriptive word embedding vectors for each descriptive word. Next, the descriptive words are numbered according to their order in the image description information. Further, based on the numbers corresponding to the target word embedding features and the descriptive word embedding features, the descriptive word embedding vectors are concatenated in ascending order of their numbers to obtain a complete descriptive word embedding sequence. Finally, the concatenated descriptive word embedding sequence is determined as the image description embedding vector.

[0218] Figure 14 illustrates the process of generating an image description vector based on image description information as a conditional constraint, and controlling image denoising based on this constraint. Specifically, the input image instance is a clock image. The image description information for the input image instance is "A Photo of clock," and "clock" in the image description information is used as the target word, represented as "S." * At this point, the image description information becomes "A Photo of S". *Based on this, the image description information is input into the word segmenter, which segments the image description information into words and determines the index of each word according to the preset index corresponding to each candidate word in the preset vocabulary index table. Specifically, the index corresponding to the description word "A" is 508, the index corresponding to the description word "Photo" is 701, the index corresponding to the description word "of" is 73, and the index corresponding to the description word "S" is... * The corresponding index is *. Further, word embedding processing is performed on the indices of each descriptor through an embedding layer, mapping the indices of each descriptor from the numerical space to the vector space, thus obtaining the word embedding features corresponding to each descriptor. The word embedding feature corresponding to the descriptor "A" is "v". 508 The word embedding feature corresponding to the descriptive word "Photo" is "v". 701 The word embedding feature corresponding to the descriptor "of" is "v". 73 ”; the descriptive word “S” * The corresponding word embedding feature is "v". * Furthermore, the word embedding features of each descriptive word are input into the text attention calculation module for attention calculation, and the text attention calculation module outputs the image description vector c. θ (y) uses the image description vector as a conditional constraint for denoising noisy image instances by the image generator, ensuring that the predicted image instances output by the image generator closely match the content represented by the image description information. Furthermore, the predicted image instances are obtained by performing four diffusion processes (four noise addition operations) on the input image instance to obtain the noisy image instance, and then performing four denoising processes on the noisy image instance by the image generator based on the image description vector.

[0219] The advantage of this embodiment is that by using fine-tuning technology to pre-set the descriptive words in the image description information to be represented by fixed characters as target words, and using a text encoder to fine-tune and search the word embedding representation of the target words, the target word embedding features corresponding to the target words are obtained. Word embedding processing is then performed on other descriptive words, converting each descriptive word into an embedding vector. This method can bind the object description content in the image description information to the image description embedding vector, which can improve the accuracy of the generated image description embedding vector, and thus improve the effectiveness of the conditional constraint content used to generate the control information of the denoising network.

[0220] In this embodiment of the disclosure, the control network includes a first control subnetwork and a second control subnetwork.

[0221] The first and second control subnetworks are neural network structures capable of cross-attention calculation. The first and second control subnetworks have the same network structure, but their network parameters differ.

[0222] The sample object feature data includes first sample object features, second sample object features, third sample object features, and fourth sample object features. The first sample object features, second sample object features, third sample object features, and fourth sample object features are obtained by extracting features from the sample object images respectively.

[0223] The first, second, third, and fourth sample object features are all used to fine-tune the denoising process of the image generation model, but the fine-tuning steps for these features differ. All four features are obtained through the aforementioned fine-tuning network. However, the network parameters of the fine-tuning network used to extract these features are different.

[0224] Referring to Figure 15, in one embodiment, step 1030 specifically includes, but is not limited to, the following steps 1510-1550:

[0225] Step 1510: Input the image description embedding vector, the first sample object features, and the contour features into the first control sub-network to generate control information and obtain the first control information;

[0226] Step 1520: Input the image description embedding vector and the second sample object features into the second control sub-network to generate control information and obtain the second control information;

[0227] Step 1530: Determine the upsampling network control information based on the features of the third and fourth sample objects;

[0228] Step 1540: Based on the first control information, the second control information, the first sample object features, and the second sample object features, determine the downsampling network control information;

[0229] Step 1550: Integrate the upsampling network control information and the downsampling network control information into denoising network control information.

[0230] Steps 1510-1550 are described in detail below.

[0231] In step 1510, the first control information is used to fine-tune the intermediate results of the denoising network of the image generation model during downsampling.

[0232] In this specific implementation, firstly, the image description embedding vector, the first sample object features, and the contour features are input into the first control sub-network. Next, the first sample object features and the contour features are concatenated to obtain a concatenated feature vector. Further, the concatenated feature vector is linearly projected through the first control sub-network to obtain a query matrix vector; and the image description embedding vector is linearly projected through the first control sub-network to obtain a key matrix vector and a value matrix vector. Further, cross-attention calculation is performed based on the query matrix vector, the key matrix vector, and the value matrix vector to obtain the attention calculation result. Finally, the attention calculation result is transformed from the numerical space to the vector space through the first control sub-network to obtain the first control information.

[0233] The specific process of calculating attention based on the query matrix vector, key matrix vector, and value matrix vector to obtain the attention calculation result can be represented by formula (1):

[0234] Where Attention(Q, K, V) is the attention calculation result; Q is the query matrix vector; K is the key matrix vector; and V is the value matrix vector. k K is the feature dimension of the key matrix vector. T It is the transpose of the key matrix vector.

[0235] In step 1520, the second control information is used to fine-tune the result generated by the denoising network of the image generation model after downsampling.

[0236] In this specific implementation, step 1520 is similar to step 1510 described above. The difference lies in that the input information of the first control sub-network in step 1510 is the image description embedding vector, the first sample object feature, and the contour feature; while the input information of the second control sub-network in step 1520 is the image description embedding vector and the second sample object feature. Since the input information is different, the resulting control information constrains different parts of the network in the denoising network. For the sake of brevity, further details will not be provided.

[0237] In step 1530, the upsampling network control information is used to impose conditional constraints on the upsampling process of the denoising network of the image generation model, so as to fine-tune and correct the upsampling process.

[0238] In the specific implementation of this embodiment, the features of the third sample object and the features of the fourth sample object are included in the same set, the object feature information in the set is regarded as a whole, and all object feature information in the set is determined as upsampling network control information.

[0239] In step 1540, the downsampling network control information is used to impose conditional constraints on the downsampling process of the denoising network of the image generation model, so as to fine-tune and correct the downsampling process.

[0240] In the specific implementation of this embodiment, the first control information, the second control information, the first sample object features, and the second sample object features are included in the same set, and all information in the set is determined as downsampling network control information.

[0241] In step 1550, the upsampling network control information and the downsampling network control information are included in the same set, the control information in the set is regarded as a whole, and all the information in the set is determined as the denoising network control information.

[0242] As shown in Figure 16, feature extraction is performed on the sample object image using fine-tuned networks with different model parameters, yielding first, second, third, and fourth sample object features, respectively. Next, the third and fourth sample object features are integrated into upsampling network control information to constrain the upsampling process. Further, the first and second sample object features, along with pre-acquired image description embedding vectors and the contour features of the reference object, are input into the control network. The control network outputs first and second control information, which are then integrated into downsampling network control information to constrain the downsampling process. Based on this, the upsampling and downsampling network control information enables full-process control of the denoising network, improving the accuracy of the denoising process.

[0243] The advantage of this embodiment is that it generates denoising network control information based on image description information, sample object features in the sample object image, and contour features of the reference object in the background template image. This makes the condition constraints of the denoising network for the image generation model not only contain a description of the image to be generated, but also contain object fine-tuning information determined based on the contour features of the reference object and the object features of the sample object. This helps to improve the comprehensiveness and accuracy of the generated denoising network control information.

[0244] Step 340 will be described in detail below.

[0245] In step 340, based on the features of the sample stitched image and the control information of the denoising network, noise prediction is performed through the image generation model to obtain the noise prediction result of the noise reference image.

[0246] In this embodiment of the disclosure, the image generation model includes a diffusion network and a denoising network.

[0247] Diffusion networks are used to progressively add noise to the features of the stitched image until the features of the input stitched image approach pure noise.

[0248] The denoising network is used to progressively denoise the latent space vectors output by the diffusion network until the image features corresponding to the spliced ​​image features of the generated samples are obtained.

[0249] Referring to Figure 17, in one embodiment, step 340 specifically includes, but is not limited to, the following steps 1710-1730:

[0250] Step 1710: Compress the features of the stitched image of the samples to obtain the compressed image features of the samples;

[0251] Step 1720: Perform diffusion processing on the sample compressed image features based on the diffusion network to obtain the sample latent space feature vector;

[0252] Step 1730: Based on the control information of the denoising network, the latent space feature vector of the sample is denoised by the denoising network to obtain the noise prediction result.

[0253] Steps 1710-1730 are described in detail below.

[0254] In step 1710, since the sample stitched image features are composed of three image features, their feature dimension is higher than the feature dimension of the input features acceptable to the diffusion network. Therefore, firstly, the sample stitched image features are compressed to reduce their dimensionality. This is done while preserving the image feature information, thus reducing the feature dimension to obtain compressed sample image features. The compressed sample image features have a feature dimension of W*H*4, which satisfies the feature dimension of the input features acceptable to the diffusion network.

[0255] In step 1720, the latent space feature vector of the sample is used to indicate the result of adding noise to the features of the sample stitched image through the diffusion network at a fixed time step.

[0256] In this specific implementation, when performing diffusion processing on the sample compressed image features based on the diffusion network, noise is successively added to the sample compressed image features through the diffusion network until the sample compressed image features approach pure noise. The diffusion process in this embodiment can be a parameterized Markov chain.

[0257] For example, a fixed time step is set to T, and the sample compressed image features are denoised T times through the forward process of the diffusion network to generate the latent space representation corresponding to the noisy reference image. The generated latent space representation is determined as the sample latent space feature vector, where T is a positive integer.

[0258] Specifically, noise is successively added to the compressed image features of the samples through the diffusion process of the diffusion network, causing the compressed image features to lose their features one by one. After T noise additions, the compressed image features of the samples become latent space representations without any features, and these latent space representations are determined as the latent space feature vectors of the samples.

[0259] It should be noted that the latent space feature vector of a sample refers to the representation of a pure noise image that lacks image features and corresponds to a noisy reference image. The form of the latent space feature vector is the same as that of the object representation; it can be a vector representation or a matrix representation, without restriction.

[0260] In step 1730, when denoising the latent space feature vectors of the samples using the denoising network, the control information of the denoising network is used as a constraint condition to remove noise from the latent space feature vectors of the samples one by one until the latent space feature vectors of the samples are restored to image features that meet the constraint requirements of the image description information. The backward process of the denoising network can also be a parameterized Markov chain.

[0261] To save space, the specific process of denoising the latent space feature vectors of the samples using a denoising network to obtain the noise prediction results in this embodiment will be described in detail below. It will not be repeated here.

[0262] Figure 18 illustrates the noise prediction process of the image generation model. Specifically, the sample stitched image features are input into the image generation model. First, the model compresses these features to create a compressed image feature with a smaller feature dimension than the original stitched image features. Next, a diffusion network performs T noise addition operations (diffusion processing) on ​​the compressed image features, transforming them into near-pure noise, resulting in a latent space feature vector. This latent space feature vector no longer reflects the image information present in the stitched image features. Further, a denoising network, based on given control information, performs T denoising operations (denoising processing) on ​​the latent space feature vector, eliminating the noise and restoring it to its original image feature form, thus obtaining the noise prediction result. This noise prediction result, to some extent, reflects the image information present in the stitched image features.

[0263] The advantage of this embodiment is that it first compresses the sample stitched image features into sample compressed image features, ensuring that the input sample compressed image features meet the input requirements of the image generation model. Next, a diffusion network is used to add noise multiple times to the sample compressed image features, thereby reducing the detail and clarity of the sample compressed image features and obtaining the sample latent space feature vector. Further, the denoising network control information is used as a constraint condition, and the noise is removed from the sample latent space feature vector through the denoising network successively until the sample latent space feature vector is restored to an image feature that meets the constraint requirements of the image description information, thus obtaining the noise prediction result. This approach introduces image description information, the contour features of the reference object, and the object features of the sample object to fine-tune the denoising process, effectively improving the model's denoising capability and thus enhancing the accuracy of the model in generating the target image.

[0264] In this embodiment of the disclosure, the denoising network can be a U-shaped network structure (U-net network), and the denoising network includes an upsampling attention network and a downsampling attention network; the denoising network control information includes upsampling network control information and downsampling network control information.

[0265] Downsampling attention networks are used to downsample the latent space feature vectors of samples to obtain more low-dimensional features.

[0266] Upsampling attention networks are used to upsample the output of an upsampling attention network, restoring the output to a denoised image feature with the same feature dimension as the latent space feature vector of the sample.

[0267] Referring to Figure 19, in one embodiment, step 1730 specifically includes, but is not limited to, the following steps 1910-1920:

[0268] Step 1910: Fuse the downsampling network control information into the first attention matrix of the downsampling attention network to update the first attention matrix, and fuse the upsampling network control information into the second attention matrix of the upsampling attention network to update the second attention matrix;

[0269] Step 1920: Denoise the latent space feature vector of the sample by using a downsampled attention network after updating the first attention matrix and an upsampled attention network after updating the second attention matrix to obtain the noise prediction result.

[0270] Steps 1910-1920 are described in detail below.

[0271] In step 1910, the first attention matrix is ​​used to perform attention calculation during the downsampling process of the denoising network to capture the correlation between different input features (sample latent space feature vector, image description embedding vector); the second attention matrix is ​​used to perform attention calculation during the upsampling process of the denoising network to capture the correlation between different input features (output of the downsampling attention network, image description embedding vector).

[0272] In this specific implementation, the downsampling attention network includes an attention downsampling module, which comprises a first attention matrix and a residual block structure. When updating the first attention matrix, the first sample object feature and the second sample object feature in the downsampling network control information are weighted and added to the output of the first attention matrix, and the first control information and the second control information in the downsampling network control information are weighted and added to the output of the attention downsampling module. Similarly, the upsampling attention network includes an attention upsampling module, which comprises a second attention matrix and a residual block structure. When updating the second attention matrix, the third sample object feature and the fourth sample object feature in the upsampling network control information are weighted and added to the output of the second attention matrix.

[0273] As shown in Figure 20, the downsampling attention network includes two attention downsampling modules, each with a first attention matrix. The features of the first sample object are weighted and added to the output of the first attention matrix; the features of the second sample object are weighted and added to the output of the second attention matrix. Simultaneously, first control information is weighted and added to the output of the first attention downsampling module; second control information is weighted and added to the output of the second attention downsampling module.

[0274] As shown in Figure 21, the upsampling attention network includes two attention upsampling modules. Each of the two attention upsampling modules has a second attention matrix. The third sample object features are weighted onto the output of the first second attention matrix; the fourth sample object features are weighted onto the output of the second second attention matrix.

[0275] In step 1920, when denoising the latent space feature vectors of the samples using a downsampling attention network after updating the first attention matrix and an upsampling attention network after updating the second attention matrix, firstly, the latent space feature vectors and image description embedding vectors are input into the denoising network of the image generation model. A linear projection is performed on the latent space feature vectors to obtain the query feature Q corresponding to the latent space feature vectors. Then, a linear projection is performed on the image description embedding vectors to obtain the key feature K and value feature V corresponding to the image description embedding vectors. Next, the first attention matrix of the first attention downsampling module of the downsampling attention network is used to perform cross-attention calculation on the query feature Q, key feature K, and value feature V to obtain the first attention calculation result. The first attention calculation result and the first sample object feature are then weighted according to a preset weight ratio to obtain the first weighted result. Further, the first weighted result is residual-processed using the residual block structure of the first attention downsampling module to obtain the first residual processing result. The first control information and the first residual processing result are then weighted according to preset weights to obtain the second weighted result. Next, the second weighted result is linearly projected to obtain a new query feature Q. Then, the first attention matrix of the second attention downsampling module of the downsampling attention network is used to perform cross-attention calculation on the key feature K, value feature V, and the new query feature Q, resulting in a second attention calculation result. This second attention calculation result is then weighted according to a preset weight ratio with the second sample object feature to obtain a third weighted result. Further, the third weighted result is residual-processed using the residual block structure of the second attention downsampling module to obtain a second residual processing result. Finally, the second control information and the second residual processing result are weighted according to preset weights to obtain a fourth weighted result, which is the output of the downsampling attention network.

[0276] Furthermore, the fourth weighted result and the image description embedding vector are input into the upsampling attention network. First, the fourth weighted result is linearly projected to obtain the query features. Then, the image description embedding vector is linearly projected to obtain the key features and value features corresponding to the image description embedding vector. Next, the second attention matrix of the first attention upsampling module of the upsampling attention network is used to perform cross-attention calculation on the key features, value features, and query features to obtain the third attention calculation result. The third attention calculation result and the third sample object features are then weighted according to a preset weight ratio to obtain the fifth weighted result. Further, the fifth weighted result is residual-processed through the residual block structure of the first attention upsampling module to obtain the third residual processing result. Finally, the third residual processing result is linearly projected to obtain new query features. The second attention matrix of the second attention upsampling module of the upsampling attention network is used to perform cross-attention calculation on the key features, value features, and new query features to obtain the fourth attention calculation result. The fourth attention calculation result and the fourth sample object features are then weighted according to a preset weight ratio to obtain the sixth weighted result. Furthermore, the residual block structure of the second attention upsampling module is used to process the residual of the sixth weighted result to obtain the denoising result of the latent space feature vector of the sample generated at prediction time step T by the denoising network; the above process is repeated to perform T-1 denoising operations on the denoising result of the latent space feature vector of the sample to obtain the noise prediction result.

[0277] The advantage of this embodiment is that it introduces corresponding denoising network control information at different stages of the denoising process to fine-tune the process. Specifically, the first and second sample object features in the downsampling network control information are fused into the first attention matrix of the downsampling attention network, and the third and fourth sample object features in the upsampling network control information are fused into the second attention matrix of the upsampling attention network. This enables the correction of the output results of multiple attention matrices in the denoising network. Furthermore, this embodiment also corrects the output results of the downsampling attention network through the first and second control information, performing multiple fine-tuning steps in each denoising process. This improves the stability and accuracy of noise prediction, making the final noise prediction results more closely match real-world needs and enhancing model training effectiveness.

[0278] Step 350 will be described in detail below.

[0279] In step 350, an image generation model is trained based on the comparison between the noise baseline image and the noise prediction results of multiple image-text sample pairs.

[0280] Referring to Figure 22, in one embodiment, step 350 specifically includes, but is not limited to, the following steps 2210-2230:

[0281] Step 2210: For each image-text sample pair, obtain the reference noise in the noise reference image, and calculate the sub-loss function of the image-text sample pair based on the comparison result of the reference noise and the noise prediction result corresponding to the noise reference image.

[0282] Step 2220: Determine the total loss function based on the sub-loss function of each image-text sample pair;

[0283] Step 2230: Train the image generation model based on the total loss function.

[0284] Steps 2210-2230 are described in detail below.

[0285] In step 2210, the reference noise is used to indicate the degree of noise added to the reference image, and the sub-loss function is used to indicate the degree of difference between the reference noise and the predicted noise of a single image-text sample pair.

[0286] In this specific implementation, firstly, for each image-text sample pair, noise is extracted from the noise reference image to obtain the reference noise in the noise reference image. Next, the noise difference between the reference noise and the noise prediction result is calculated to obtain the noise difference result. Finally, a sub-loss function is calculated based on the noise difference result.

[0287] In step 2220, the total loss function is used to indicate the overall difference between the reference noise and the predicted noise of all image-text sample pairs. The smaller the total loss function, the smaller the overall difference between the reference noise and the predicted noise of all image-text sample pairs, and the higher the image generation accuracy of the image generation model.

[0288] In this specific implementation, the sub-loss functions of each image-text sample pair are averaged to obtain the total loss function. Specifically, first, the total number of image-text sample pairs is determined. Next, the sub-loss functions of all image-text sample pairs are summed to obtain the sum of the sub-loss functions. Finally, the total loss function is obtained by dividing the sum of the sub-loss functions by the total number of image-text sample pairs.

[0289] In step 2230, with the goal of minimizing the total loss function, the model parameters of the image generation model are adjusted, and steps 310-350 above are repeated to achieve iterative training of the image generation model. The model parameters that minimize the total loss function are taken as the final model parameters, and the image generation model with the final model parameters is taken as the trained image generation model.

[0290] The advantage of this embodiment is that, based on supervised learning, it determines the sub-loss function of each image-text sample pair according to the noise difference between the reference noise and the noise prediction result of each image-text sample pair, and constructs the total loss function based on multiple sub-loss functions. This approach can train the image generation model by minimizing the difference between the reference noise and the noise prediction result, which is beneficial to improving the model training effect and thus improving the model's prediction accuracy of the generated image.

[0291] In this embodiment of the disclosure, the noise prediction result is obtained through prediction at multiple prediction time steps.

[0292] The prediction time step is used to indicate the frequency of adding noise and denoising the sample stitched image features of the model generated from the input image.

[0293] Referring to Figure 23, in one embodiment, step 2210 specifically includes, but is not limited to, the following steps 2310-2330:

[0294] Step 2310: Based on the noise prediction results, determine the prediction noise for the last prediction time step;

[0295] Step 2320: Calculate the regularization term based on the reference noise and the predicted noise to obtain the regularization term calculation result;

[0296] Step 2330: Determine the sub-loss function based on the calculation results of the regularization term.

[0297] Steps 2310-2330 are described in detail below.

[0298] In step 2310, the predicted noise is used to indicate the noise contained in the denoised result (denoised image) produced by the image generation model at the last prediction time step.

[0299] In this specific implementation, the image generation model denoises the denoised result (denoised image) generated in the previous prediction time step again at each prediction time step, and after multiple prediction time steps of denoising processing, a noise prediction result is obtained. The noise contained in the denoised results of different prediction time steps is different. Based on this, the noise prediction result includes the denoised result (denoised image) generated in the last prediction time step. The predicted noise of the last prediction time step can be directly extracted from the noise prediction result.

[0300] In step 2320, the result of the regularization term calculation is used to indicate the degree of difference between the reference noise and the predicted noise.

[0301] In the specific implementation of this embodiment, firstly, for each image-text sample pair, the difference between the reference noise and the predicted noise is calculated to obtain the noise difference. Next, the noise difference is regularized to obtain the regularization term calculation result.

[0302] In step 2330, the regularization term calculation result is transformed based on the denoising network control information, converting the regularization term calculation result into the form of a conditional loss function, and the transformed loss function is determined as a sub-loss function. The sub-loss function in this embodiment can be expressed as shown in formula (2):

[0303] Among them, L LDM It is the sub-loss function; ε is the reference noise. The reference noise is indicated to follow a standard normal distribution (mean 0, variance 1). t represents the prediction time step, which is used to progressively update the latent space feature vectors of the samples during image generation. t z represents the latent space feature vector of the sample at prediction time step t. t This is the result of adding noise to the sample compressed image feature z at prediction time step t. y and c both refer to the control information of the denoising network, used to impose conditional constraints on the image generation process. x is the noisy baseline image used for training. Used to indicate the encoder's expectation of the input x with respect to ε. ∈ θ This refers to the process of analyzing the latent space feature vector z of a sample at prediction time step t, based on the conditional constraint c. t Predicted noise after denoising. This is the result of the regularization term calculation.

[0304] The advantage of this embodiment is that by calculating the regularization loss between the reference noise and the noise prediction result of the image-text sample pair, the regularization calculation result of each image-text sample pair is obtained. The regularization calculation result (L2 loss) is used as a sub-loss function to train the image generation model. This allows the denoising network control information to better adjust the image generation process of the image generation model, making the generation process more controllable. At the same time, the addition of conditional information (denoising network control information) to the sub-loss function can closely correlate the generated noise prediction result with the given conditions, improving the consistency and accuracy of the generated image, and thus improving the image quality of the target image generated by the model.

[0305] Figure 24 shows the overall flowchart of model training according to an embodiment of this disclosure. Specifically, the background template image, the expected result image, and the image description information of the expected result image are used as input.

[0306] First, a reference noise is generated using a random number i, and this reference noise is added to the expected effect image to obtain a noise baseline image. The specific process is similar to steps 510-530 above. Next, the background template image is masked to obtain an object replacement mask (template mask image). The specific process is similar to steps 810-820 above.

[0307] Furthermore, the template mask image, the noisy reference image, and the background template image are encoded separately using an encoder, and the three encoded results are concatenated to form the sample stitched image features. The sample stitched image features are compressed into sample compressed image features Z using an image generation model, and then diffused through a diffusion network. Noise is added to the sample compressed image features Z T times to obtain the sample latent space feature vector Z. T Furthermore, the image description information of the expected effect image is used as a conditional constraint and converted into a text form that meets the input requirements of the text encoder τ. The image description information of the expected effect image is then input into the text encoder, which outputs an image description embedding vector. The specific process is similar to steps 1310-1340 above. Further, the object bounding boxes corresponding to the background template image are extracted, and the outline line drawing (outline feature) of the reference object in the background template image is generated based on the extracted object bounding boxes. The specific process is similar to steps 1110-1130 above. Simultaneously, based on the sample object image of the sample object, sample object feature data is generated, wherein the sample object feature data includes a first object feature QKV1-A*, a second object feature QKV2-A*, a third object feature QKV3-A*, and a fourth object feature QKV4-A*. Furthermore, the first object feature QKV1-A*, the second object feature QKV2-A*, the contour feature, and the image description embedding vector are input into the control network. The first control subnetwork of the control network outputs the first control information, and the second control subnetwork of the control network outputs the second control information.

[0308] Next, the denoising network processes the latent space feature vector Z of the samples. T During denoising, based on the first control information, the first object feature QKV1-A*, and the second object feature QKV2-A*, the latent space feature vector Z of the sample is first processed through the attention matrices of the downsampling attention network of the denoising network. T Downsampling is performed to obtain the downsampling result, and the downsampling result is corrected using the second control information to obtain the corrected downsampling result. Then, based on the third object feature QKV3-A* and the fourth object feature QKV4-A*, the corrected downsampling result is upsampled through the attention matrices of the upsampling attention network of the denoising network to obtain the denoising result Z at the prediction time step T. T-1' Furthermore, a denoising network is used to refine the denoising result Z at prediction time step T.T-1' Following the above process for denoising, after T-1 denoising iterations, the denoised result Z is obtained. ’ The specific process is similar to the steps mentioned above from 1910 to 1920.

[0309] Finally, based on the denoising result Z ’ Noise prediction is performed to obtain the noise prediction result. A loss function LOSS is constructed based on the reference noise and the noise prediction result. The image generation model is then iteratively trained based on the loss function LOSS until the image generation model meets the training requirements. The specific process is similar to step 350 above. To save space, it will not be described in detail.

[0310] The image generation method of one embodiment of this disclosure will now be described in detail.

[0311] According to one embodiment of this disclosure, an image generation method is provided.

[0312] This image generation method is generally applied in business scenarios where objects in a fixed background image need to be replaced with target people, target items, etc., such as the video production scenario and item display scenario shown in Figures 2A-2C. This disclosure provides a scheme for image generation based on image description information and object information (the outline information of the reference object in the background image and the object features of the target object to be replaced), using an image generation model, which can improve the accuracy of target image generation.

[0313] As shown in Figure 25, an image generation method according to an embodiment of this disclosure can be executed by an electronic device, which may be the image processing server or object terminal shown in Figure 1. The image generation method may include:

[0314] Step 2510: Obtain the object image containing the first object, the background image, and the description information;

[0315] Step 2520: Determine the features of the stitched image based on the background image, the preset noise image, and the mask image;

[0316] Step 2530: Based on the contour features, description information, and object image of the second object in the background image, determine the denoising control information of the image generation model;

[0317] Step 2540: Based on the features of the stitched image and the noise reduction control information, the image is generated through an image generation model to obtain the target image.

[0318] Steps 2510-2540 are described in detail below.

[0319] In step 2510, an object image containing the first object, a background image, and descriptive information are obtained.

[0320] An object image containing the first object refers to an image that reflects the characteristics of the first object. Specifically, one or more object images containing the first object can be obtained.

[0321] The background image refers to the image that provides background content for the final image to be generated, that is, the background image to which the first object is to be added. The background image contains the second object that will be replaced by the first object in the object image.

[0322] The description information is used to describe the replacement from the second object to the first object.

[0323] In this specific implementation, step 2510 is similar to the process described in step 310 above for obtaining the sample object image, background template image, and image description information. To save space, it will not be repeated here.

[0324] In step 2520, the features of the stitched image are determined based on the background image, the preset noise image, and the mask image.

[0325] The preset noise image is a Gaussian-distributed noise image generated based on random numbers, and its specific generation method is similar to step 520 above. The difference is that the random numbers used to generate the preset noise image in step 2520 are different from the random numbers in step 520. To save space, this will not be elaborated further.

[0326] The masked image is obtained by masking the second object in the background image, and the specific process is similar to steps 810-820 above. To save space, it will not be described in detail again.

[0327] The stitched image features are used to indicate the stitched result of features from the background image, the preset noise image, and the mask image.

[0328] To save space, the specific process of determining the features of the target stitched image based on the background image, the preset noise image, and the mask image in this embodiment will be described in detail below. It will not be repeated here.

[0329] In step 2530, based on the contour features, description information, and object image of the second object in the background image, the denoising control information of the image generation model is determined.

[0330] The image generation model is generated according to the training method of the image generation model in the above embodiment.

[0331] Denoising control information is used as a conditional constraint to assist the image generation model in denoising the image during image generation, so as to improve the accuracy of image denoising and make the image denoising effect closer to the real requirements.

[0332] In the specific implementation of this embodiment, the specific process of step 2530 is similar to that of step 330 described above. To save space, it will not be described again.

[0333] In step 2540, based on the features of the stitched image and the noise reduction control information, the image is generated by the image generation model to obtain the target image.

[0334] The target image is used to indicate the result of replacing the second object in the background image with the first object in the target image.

[0335] To save space, the specific process of generating the target image based on the stitched image features and denoising control information through an image generation model, as described in this embodiment, will be described in detail below. It will not be repeated here.

[0336] Through steps 2510-2540, in this embodiment of the disclosure, when generating an image using an image generation model, the image features of the background image, the preset noise image, and the mask image are integrated into stitched image features. This allows the stitched image features to contain background image information from the background image and information about the second object, while also incorporating noise information, making the stitched image features more realistic. Furthermore, the contour features and descriptive information of the second object in the background image, along with the first object image, are introduced to generate denoising control information for the denoising network of the image generation model. This enables the image generation model to fine-tune objects during the denoising process, ensuring that the objects in the generated target image meet the constraints of the descriptive information, the object characteristics of the first object, and the contour features of the second object. This results in better consistency and coordination between the first object and the background in the target image, thereby improving the accuracy of the model in generating the target image.

[0337] Referring to Figure 26, in one embodiment, step 2520 specifically includes, but is not limited to, the following steps 2610-2640:

[0338] Step 2610: Perform a first encoding process on the preset noisy image to obtain the noisy image features;

[0339] Step 2620: Perform a second encoding process on the background image to obtain background image features;

[0340] Step 2630: Perform a third encoding process on the mask image to obtain the mask image features;

[0341] Step 2640: Concatenate the noise image features, background image features, and mask image features to obtain the concatenated image features.

[0342] Steps 2610-2640 are described in detail below.

[0343] Noise image features are used to indicate the result of converting a preset noisy image from data space to pixel space.

[0344] Background image features are used to indicate the result of converting the background image from data space to pixel space.

[0345] Mask image features are used to indicate the result of converting a mask image from data space to pixel space.

[0346] In the specific implementation of this embodiment, the specific process of steps 2610-2640 is similar to that of steps 710-740 described above. To save space, it will not be described again.

[0347] The advantage of this embodiment is that it transforms the image information of the preset noisy image, background image, and mask image from pixel space to latent vector space, and then concatenates the image information (noise image features, background image features, and mask image features) of the preset noisy image, background image, and mask image in the latent vector space to form a concatenated image feature with multiple image information. Furthermore, using the concatenated image feature as input to the image generation model during image generation allows the model's input data to incorporate noise information, background information, and information about a second object in the background. This significantly improves the richness and comprehensiveness of the feature information in the concatenated image feature, which is beneficial for improving the accuracy of the model's image generation and making the final generated target image more realistic and accurate.

[0348] In this embodiment of the disclosure, the image generation model includes a diffusion network, a denoising network, and a decoding network.

[0349] The decoding network is used to transform the image features in the denoising result from the latent vector space to the pixel space.

[0350] Referring to Figure 27, in one embodiment, step 2540 specifically includes, but is not limited to, the following steps 27710-2740:

[0351] Step 2710: Compress the features of the stitched image to obtain compressed image features;

[0352] Step 2720: Perform diffusion processing on the compressed image features based on the diffusion network to obtain the latent space feature vector;

[0353] Step 2730: Based on the denoising control information, the latent space feature vector is denoised through a denoising network to obtain the denoising result;

[0354] Step 2740: Perform feature decoding on the denoising result based on the decoding network to obtain the target image.

[0355] Steps 2710-2740 are described in detail below.

[0356] Compressed image features are used to indicate the dimensionality reduction result of stitched image features. The compressed image features and the stitched image features have the same image feature information; however, the feature dimension of the compressed image features is lower than that of the stitched image features.

[0357] The latent space feature vector is used to indicate the result of adding noise to the compressed image features through a diffusion network at a fixed time step.

[0358] The denoising result is used to indicate the image features that meet the constraints of image description information generated by the denoising network denoising the latent space feature vector at a fixed time step.

[0359] In the specific implementation of this embodiment, the specific process of steps 2710-2730 is similar to that of steps 1710-1730 described above. To save space, it will not be described again.

[0360] In step 2740, the denoising result is input into the decoding network. The decoding network maps the image features in the denoising result from the latent vector space back to the original pixel space, generating a noise-free target image that matches the description information.

[0361] The advantage of this embodiment is that it first compresses the stitched image features into compressed image features, ensuring that the input compressed image features meet the input requirements of the image generation model. Next, a diffusion network is used to add noise multiple times to the compressed image features, reducing their detail and clarity to obtain a latent space feature vector. Further, a denoising network is used to denoise the latent space feature vector multiple times, enabling prediction of the source image. During this process, denoising control information is used to fine-tune the predicted image features, making the object features in the final denoised result more closely match the real-world requirements. Finally, a decoding network is used to decode the denoised result, thereby obtaining the predicted target image. This improves the image quality of the target image, resulting in better harmony between the object and the background.

[0362] Figure 28 illustrates a specific application module of the image generation method according to an embodiment of this disclosure. Specifically, firstly, for a background image containing a second object, the background image undergoes object face masking processing to obtain a mask image; the specific process is similar to steps 810-820 above. Key points of the object masking area are then extracted from the background image to obtain the contour features of the second object in the background image; the specific process is similar to steps 1110-1130 above. Furthermore, for the first object to replace the second object, feature extraction is performed on the object image including the first object to obtain object fine-tuning feature data of the first object. Further, a latent space input is constructed based on the random noise image, background image, and mask image to obtain stitched image features; the specific process is similar to steps 2610-2640 above. Next, the contour features and descriptive information of the target image to be generated are used to drive the control network. The control network outputs denoising control information as an auxiliary (conditional constraint), and the image generation model performs diffusion processing and denoising processing on the stitched image features to obtain a denoising result. Finally, the denoising result is converted into the target image and output using a decoding network, thereby generating a target image with the background image as the background and the first object as the foreground. The specific process is similar to steps 2710-2740 above. To save space, it will not be described in detail again.

[0363] Figure 29 shows an overall flowchart of image generation based on an image generation model according to an embodiment of this disclosure. Specifically, the background image and the descriptive information corresponding to the image to be generated are taken as inputs.

[0364] First, a random noise map is generated using a random number i. Next, the background image is masked to obtain an object replacement mask (mask image), the specific process of which is similar to steps 810-820 above. Further, the random noise map, mask image, and background image are encoded separately using an encoder, and the three encoding results are concatenated to form a stitched image feature.

[0365] Furthermore, the stitched image features are compressed into compressed image features Z using an image generation model, and then the compressed image features Z are diffused using a diffusion network. Noise is added to the compressed image features Z T times to obtain the latent space feature vector Z. TFurther, the descriptive information corresponding to the image to be generated is used as a conditional constraint and converted into a text form that meets the input requirements of the text encoder τ. The descriptive information is then input into the text encoder, which outputs an image description embedding vector. The specific process is similar to steps 1310-1340 above. Further, the object bounding box corresponding to the background image is extracted, and the outline line drawing (outline feature) of the second object in the background image is generated based on the extracted object bounding box. The specific process is similar to steps 1110-1130 above. Simultaneously, object feature data is generated based on the object image including the first object. The object feature data includes the first object feature QKV1-A*, the second object feature QKV2-A*, the third object feature QKV3-A*, and the fourth object feature QKV4-A*. Further, the first object feature QKV1-A*, the second object feature QKV2-A*, the outline feature, and the image description embedding vector are input into the control network. The first control subnetwork of the control network outputs first control information, and the second control subnetwork of the control network outputs second control information.

[0366] Next, the denoising network processes the latent space feature vector Z. T During denoising, based on the first control information, the first object feature QKV1-A*, and the second object feature QKV2-A*, the latent space feature vector Z is first processed by the downsampling attention network of the denoising network through the attention matrices of each attention matrix. T Downsampling is performed to obtain the downsampling result, and the downsampling result is corrected using the second control information to obtain the corrected downsampling result. Then, based on the third object feature QKV3-A* and the fourth object feature QKV4-A*, the corrected downsampling result is upsampled through the attention matrices of the upsampling attention network of the denoising network to obtain the denoising result Z at the prediction time step T. T-1' Furthermore, a denoising network is used to refine the denoising result Z at prediction time step T. T-1' Following the above process for denoising, after T-1 denoising iterations, the denoised result Z is obtained. ’ The specific process is similar to steps 1910-1920 described above. Finally, based on the decoding network, the denoising result Z is... ’ Feature decoding is performed to obtain the target image I, and the specific process is similar to step 2740 above. To save space, it will not be described in detail again.

[0367] The apparatus and device according to embodiments of this disclosure will now be described.

[0368] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0369] It should be noted that in various specific embodiments of this application, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require obtaining target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this application be obtained.

[0370] Figure 30 is a schematic diagram of the structure of a training device 3000 for an image generation model provided in an embodiment of this disclosure. The training device 3000 for the image generation model includes:

[0371] The first acquisition unit 3010 is used to acquire multiple image-text sample pairs, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including the sample object. The noise reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object.

[0372] The first determining unit 3020 is used to determine the features of the sample stitched image based on the noise reference image, the background template image, and the template mask image, wherein the template mask image is obtained by masking the reference object in the background template image;

[0373] The second determining unit 3030 is used to determine the denoising network control information of the image generation model based on the contour features of the reference object in the background template image, the image description information, and the sample object image.

[0374] The prediction unit 3040 is used to predict noise based on the features of the sample stitched image and the control information of the denoising network, and to obtain the noise prediction result corresponding to the noise reference image through the image generation model.

[0375] Training unit 3050 is used to train the image generation model based on the comparison results of the noisy baseline image and its corresponding noise prediction results in multiple image-text sample pairs.

[0376] Optionally, the training unit 3050 includes:

[0377] The calculation module (not shown) is used to obtain the reference noise in the noise reference image for each image-text sample pair, and calculate the sub-loss function of the image-text sample pair based on the comparison result of the reference noise and the noise prediction result corresponding to the noise reference image. The reference noise is the noise added to the reference image.

[0378] A determination module (not shown) is used to determine the total loss function based on the sub-loss function of each image-text sample pair;

[0379] A training module (not shown) is used to train the image generation model based on the total loss function.

[0380] Optionally, the noise prediction result is obtained by prediction over multiple prediction time steps;

[0381] The calculation module (not shown) is used for:

[0382] Based on the noise prediction results, determine the prediction noise for the last prediction time step;

[0383] The regularization term is calculated based on the reference noise and the predicted noise, and the calculation result of the regularization term is obtained.

[0384] Based on the calculation results of the regularization term, the sub-loss function is determined.

[0385] Optionally, the first determining unit 3020 is used for:

[0386] The noisy reference image is subjected to a first encoding process to obtain the reference image encoding features;

[0387] The background template image is subjected to a second encoding process to obtain the background image encoding features;

[0388] A third encoding process is performed on the template mask image to obtain the mask image encoding features;

[0389] The coded features of the reference image, the coded features of the background image, and the coded features of the mask image are concatenated to obtain the coded features of the sample image.

[0390] Optionally, the second determining unit 3030 includes:

[0391] An encoding module (not shown) is used to encode the image description information to obtain an image description embedding vector;

[0392] An extraction module (not shown) is used to extract features from the sample object image to obtain sample object feature data;

[0393] The generation module (not shown) is used to generate control information through a preset control network based on the image description embedding vector, sample object feature data, and the contour features of the reference object, thereby obtaining the denoising network control information.

[0394] Optionally, the encoding module (not shown) is used for:

[0395] The image description information is segmented into multiple descriptive words;

[0396] The target word is identified from multiple descriptive words, and the target word embedding feature corresponding to the target word is found based on a pre-defined dictionary;

[0397] For each other descriptive word besides the target word, word embedding is performed on each other descriptive word to obtain the descriptive word embedding features of each other descriptive word.

[0398] The target word embedding features and the embedding features of each descriptive word are integrated into an image description embedding vector.

[0399] Optionally, the control network includes a first control sub-network and a second control sub-network; the sample object feature data includes a first sample object feature, a second sample object feature, a third sample object feature, and a fourth sample object feature, wherein the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained by extracting features from the sample object image respectively;

[0400] The generation module (not shown) is used for:

[0401] The image description embedding vector, the first sample object features, and the contour features are input into the first control sub-network to generate control information, thereby obtaining the first control information;

[0402] The image description embedding vector and the features of the second sample object are input into the second control sub-network to generate control information, thus obtaining the second control information.

[0403] Based on the features of the third and fourth sample objects, the control information of the upsampling network is determined.

[0404] Based on the first control information, the second control information, the characteristics of the first sample object, and the characteristics of the second sample object, the control information of the downsampling network is determined;

[0405] The upsampling network control information and the downsampling network control information are integrated into the denoising network control information.

[0406] Optionally, the noisy reference image is generated in the following way:

[0407] The background template image will be replaced with a baseline image of the sample object;

[0408] Based on a predetermined random number generation model, random numbers that follow a Gaussian distribution are generated;

[0409] For each pixel in the reference image, a random number is added to the pixel value of that pixel to obtain a noisy reference image.

[0410] Optionally, the image generation model includes a diffusion network and a denoising network;

[0411] Prediction unit 3040 includes:

[0412] A compression module (not shown) is used to compress the features of the sample stitched image to obtain the compressed image features of the sample.

[0413] A diffusion module (not shown) is used to perform diffusion processing on the sample compressed image features based on a diffusion network to obtain the sample latent space feature vector;

[0414] A denoising module (not shown) is used to denoise the latent space feature vector of the sample through the denoising network based on the denoising network control information to obtain the noise prediction result.

[0415] Optionally, the denoising network includes an upsampling attention network and a downsampling attention network; the denoising network control information includes upsampling network control information and downsampling network control information.

[0416] The noise reduction module (not shown) is used for:

[0417] The control information of the downsampled network is fused into the first attention matrix of the downsampled attention network to update the first attention matrix, and the control information of the upsampled network is fused into the second attention matrix of the upsampled attention network to update the second attention matrix.

[0418] By using a downsampled attention network after updating the first attention matrix and an upsampled attention network after updating the second attention matrix, the latent space feature vectors of the samples are denoised to obtain noise prediction results.

[0419] Optionally, the sample object image is generated in the following way:

[0420] Obtain a sample image containing sample objects;

[0421] Based on a preset object segmentation model, the sample image is segmented to obtain a sample segmentation image containing the sample object.

[0422] Image enhancement is performed on the segmented sample image to obtain the sample object image.

[0423] Optionally, the template mask image is generated in the following way:

[0424] Determine the object outline region of the reference object in the background template image;

[0425] In the background template image, the pixel values ​​of each pixel within the object outline area are replaced with the first value, and the pixel values ​​of each pixel outside the object outline area are replaced with the second value to obtain the template mask image.

[0426] Optionally, the contour features of the reference object in the background template image are determined in the following way:

[0427] Object detection is performed on the background template image to obtain the object skeleton map of the reference object;

[0428] Pose feature extraction is performed on the object skeleton diagram to obtain multiple object pose key points;

[0429] Contour features are determined based on multiple object pose key points.

[0430] Figure 31 is a schematic diagram of the structure of an image generation apparatus 3100 provided in an embodiment of this disclosure. The image generation apparatus 3100 includes:

[0431] The second acquisition unit 3110 is used to acquire an object image including a first object, a background image, and descriptive information, wherein the background image contains the second object, and the descriptive information is used to describe the replacement from the second object to the first object;

[0432] The third determining unit 3120 is used to determine the features of the spliced ​​image based on the background image, the preset noise image, and the mask image, wherein the mask image is obtained by masking the second object in the background image;

[0433] The fourth determining unit 3130 is used to determine the denoising control information of the image generation model based on the contour features, description information, and object image of the second object in the background image. The image generation model is generated by the training method of the image generation model described above.

[0434] The image generation unit 3140 is used to generate an image based on the features of the stitched image and the noise reduction control information through an image generation model to obtain a target image, wherein the target image is used to indicate the result of replacing a second object in the background image with a first object in the target image.

[0435] Optionally, the third determining unit 3120 is used for:

[0436] The preset noisy image is subjected to a first encoding process to obtain the noisy image features;

[0437] The background image undergoes a second encoding process to obtain its features;

[0438] A third encoding process is performed on the mask image to obtain the mask image features;

[0439] The features of the noisy image, the background image, and the mask image are concatenated to obtain the concatenated image features.

[0440] Optionally, the image generation model includes a diffusion network, a denoising network, and a decoding network;

[0441] Image generation unit 3140 is used for:

[0442] The features of the stitched image are compressed to obtain compressed image features;

[0443] The features of compressed images are diffused using a diffusion network to obtain latent space feature vectors.

[0444] Based on the denoising control information, the latent space feature vector is denoised through a denoising network to obtain the denoising result;

[0445] The denoising result is feature-decoded using a decoding network to obtain the target image.

[0446] Referring to Figure 32, which is a structural block diagram of a terminal implementing the image generation model training method or image generation method of the present disclosure, the terminal may be the terminal shown in Figure 1. The terminal includes components such as: a radio frequency (RF) circuit 3210, a memory 3215, an input unit 3230, a display unit 3240, a sensor 3250, an audio circuit 3260, a wireless fidelity (WiFi) module 3270, a processor 3280, and a power supply 3290. Those skilled in the art will understand that the terminal structure shown in Figure 32 does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0447] The RF circuit 3210 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 3280; in addition, it transmits uplink data to the base station.

[0448] The memory 3215 can be used to store software programs and modules. The processor 3280 executes various functional applications and data processing of the target terminal by running the software programs and modules stored in the memory 3215.

[0449] The input unit 3230 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the target terminal. Specifically, the input unit 3230 may include a touch panel 3231 and other input devices 3232.

[0450] Display unit 3240 can be used to display input or provided information, as well as various menus of the target terminal. Display unit 3240 may include display panel 3241.

[0451] Audio circuitry 3260, speaker 3261, and microphone 3262 provide an audio interface.

[0452] In this embodiment, the processor 3280 included in the terminal can execute the training method or image generation method of the image generation model in the previous embodiment.

[0453] Figure 33 is a structural block diagram of a server portion of the training method or image generation method for implementing the image generation model of the present disclosure. The server can be the image processing server shown in Figure 1. Servers can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 3322 (e.g., one or more processors) and memory 3332, and one or more storage media 3330 (e.g., one or more mass storage devices) for storing application programs 3342 or data 3344. The memory 3332 and storage media 3330 can be temporary or persistent storage. The program stored in the storage media 3330 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the server. Furthermore, the CPU 3322 may be configured to communicate with the storage media 3330 and execute the series of instruction operations in the storage media 3330 on the server.

[0454] The server may also include one or more power supplies 3326, one or more wired or wireless network interfaces 3350, one or more input / output interfaces 3358, and / or one or more operating systems 3341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0455] The central processing unit 3322 in the server can be used to execute the training method or image generation method of the image generation model of the present disclosure embodiments.

[0456] This disclosure also provides a computer-readable storage medium for storing a computer program for executing the training method or image generation method of the image generation model in the foregoing embodiments.

[0457] This disclosure also provides a computer program product, which includes a computer program. An electronic device's processor reads and executes the computer program, causing the electronic device to perform the training method or image generation method that implements the image generation model described above.

[0458] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0459] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0460] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0461] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0462] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0463] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0464] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0465] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0466] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A method for training an image generation model, executed by an electronic device, the training method comprising: Multiple image-text sample pairs are obtained, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including the sample object. The noise reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object. Based on the noise reference image, the background template image, and the template mask image, the features of the sample stitched image are determined, wherein the template mask image is obtained by masking the reference object in the background template image; Based on the contour features of the reference object in the background template image, the image description information, and the sample object image, the denoising network control information of the image generation model is determined; Based on the features of the sample stitched image and the control information of the denoising network, noise prediction is performed through the image generation model to obtain the noise prediction result corresponding to the noise reference image; The image generation model is trained based on the comparison results between the noise reference image and its corresponding noise prediction result in multiple image-text sample pairs.

2. The training method for the image generation model according to claim 1, wherein, The step of training the image generation model based on the comparison results between the noise reference image and its corresponding noise prediction result in multiple image-text sample pairs includes: For each image-text sample pair, a reference noise in the noise reference image is obtained, and a sub-loss function for the image-text sample pair is calculated based on the comparison result between the reference noise and the noise prediction result corresponding to the noise reference image. The reference noise is noise added to the reference image. The total loss function is determined based on the sub-loss function for each of the image-text sample pairs; The image generation model is trained based on the total loss function.

3. The training method for the image generation model according to claim 2, wherein, The noise prediction result is obtained through prediction at multiple prediction time steps; The calculation of the sub-loss function for the image-text sample pair based on the comparison result between the reference noise and the noise baseline image includes: Based on the noise prediction results, the prediction noise for the last prediction time step is determined; The regularization term is calculated based on the reference noise and the predicted noise to obtain the regularization term calculation result. Based on the calculation results of the regularization term, the sub-loss function is determined.

4. The training method for the image generation model according to any one of claims 1 to 3, wherein, The step of determining the features of the sample stitched image based on the noise reference image, the background template image, and the template mask image includes: The noisy reference image is subjected to a first encoding process to obtain the reference image encoding features; The background template image is subjected to a second encoding process to obtain the background image encoding features; The template mask image is subjected to a third encoding process to obtain the mask image encoding features; The reference image coding features, the background image coding features, and the mask image coding features are concatenated to obtain the sample concatenated image features.

5. The training method for the image generation model according to any one of claims 1 to 4, wherein, The step of determining the denoising network control information of the image generation model based on the contour features of the reference object in the background template image, the image description information, and the sample object image includes: The image description information is encoded to obtain an image description embedding vector; Feature extraction is performed on the image of the sample object to obtain the feature data of the sample object; Based on the image description embedding vector, the sample object feature data, and the contour features of the reference object, control information is generated through a preset control network to obtain the denoising network control information.

6. The training method for the image generation model according to claim 5, wherein, The process of encoding the image description information to obtain an image description embedding vector includes: The image description information is segmented into words to obtain multiple descriptive words; The target word is identified from among the multiple descriptive words, and the target word embedding feature corresponding to the target word is found based on a preset dictionary; For each of the other descriptive words besides the target word, word embedding processing is performed on each of the other descriptive words to obtain the descriptive word embedding features of each of the other descriptive words. The target word embedding features and each of the descriptive word embedding features are integrated into the image description embedding vector.

7. The training method for the image generation model according to claim 5 or 6, wherein, The control network includes a first control sub-network and a second control sub-network; the sample object feature data includes a first sample object feature, a second sample object feature, a third sample object feature, and a fourth sample object feature, wherein the first sample object feature, the second sample object feature, the third sample object feature, and the fourth sample object feature are obtained by extracting features from the sample object image respectively; The process of generating control information for the denoising network based on the image description embedding vector, the sample object feature data, and the contour features of the reference object through a preset control network, includes: The image description embedding vector, the first sample object features, and the contour features are input into the first control sub-network to generate control information, thereby obtaining the first control information; The image description embedding vector and the features of the second sample object are input into the second control sub-network to generate control information, thereby obtaining the second control information. Based on the features of the third sample object and the features of the fourth sample object, the upsampling network control information is determined; Based on the first control information, the second control information, the first sample object features, and the second sample object features, the downsampling network control information is determined; The upsampling network control information and the downsampling network control information are integrated into the denoising network control information.

8. The training method for the image generation model according to any one of claims 1 to 7, wherein, The noise reference image is generated in the following manner: The reference object in the background template image is determined to be replaced with a baseline image of the sample object; Based on a predetermined random number generation model, random numbers that follow a Gaussian distribution are generated; For each pixel in the reference image, the random number is added to the pixel value of the pixel to obtain the noise reference image.

9. The training method for the image generation model according to any one of claims 1 to 8, wherein, The image generation model includes a diffusion network and a denoising network; The step of performing noise prediction based on the features of the sample stitched image and the control information of the denoising network, through the image generation model, to obtain the noise prediction result corresponding to the noise reference image, includes: The features of the stitched image of the sample are compressed to obtain the compressed image features of the sample. Based on the diffusion network, the sample compressed image features are diffused to obtain the sample latent space feature vector; Based on the denoising network control information, the sample latent space feature vector is denoised through the denoising network to obtain the noise prediction result.

10. The training method for the image generation model according to claim 9, wherein, The denoising network includes an upsampling attention network and a downsampling attention network; the denoising network control information includes upsampling network control information and downsampling network control information. The step of denoising the sample latent space feature vectors based on the denoising network control information to obtain the noise prediction result includes: The downsampling network control information is fused into the first attention matrix of the downsampling attention network to update the first attention matrix, and the upsampling network control information is fused into the second attention matrix of the upsampling attention network to update the second attention matrix. The noise prediction result is obtained by denoising the latent space feature vector of the sample through the downsampled attention network after updating the first attention matrix and the upsampled attention network after updating the second attention matrix.

11. The training method for the image generation model according to any one of claims 1 to 10, wherein, The sample object image is generated in the following way: Obtain a sample image containing the sample object; Based on a preset object segmentation model, the sample image is segmented to obtain a sample segmentation image containing the sample object; Image enhancement is performed on the segmented sample image to obtain the sample object image.

12. The training method for the image generation model according to any one of claims 1 to 11, wherein, The template mask image is generated in the following way: Determine the object outline region of the reference object in the background template image; In the background template image, the pixel values ​​of each pixel within the object outline area are replaced with a first value, and the pixel values ​​of each pixel outside the object outline area are replaced with a second value to obtain the template mask image.

13. The training method for the image generation model according to any one of claims 1 to 12, wherein, The contour features of the reference object in the background template image are determined in the following way: Object detection is performed on the background template image to obtain the object skeleton map of the reference object; Pose feature extraction is performed on the object skeleton diagram to obtain multiple object pose key points; The contour features are determined based on the multiple object pose key points.

14. An image generation method, performed by an electronic device, the image generation method comprising: Acquire an object image containing a first object, a background image, and descriptive information, wherein the background image contains a second object, and the descriptive information is used to describe the replacement from the second object to the first object; Based on the background image, the preset noise image, and the mask image, the features of the stitched image are determined, wherein the mask image is obtained by masking the second object in the background image; Based on the contour features of the second object in the background image, the description information, and the object image, the denoising control information of the image generation model is determined, and the image generation model is generated according to the training method of the image generation model according to any one of claims 1 to 13; Based on the stitched image features and the denoising control information, an image is generated using the image generation model to obtain a target image, wherein the target image is used to indicate the result of replacing the second object in the background image with the first object in the target image.

15. The image generation method according to claim 14, wherein, The image generation model includes a diffusion network, a denoising network, and a decoding network; The step of generating an image based on the stitched image features and the denoising control information, using the image generation model to obtain the target image, includes: The stitched image features are compressed to obtain compressed image features; The compressed image features are diffused based on the diffusion network to obtain the latent space feature vector. Based on the denoising control information, the latent space feature vector is denoised through the denoising network to obtain the denoising result. The denoising result is feature-decoded based on the decoding network to obtain the target image.

16. A training device for an image generation model, wherein, The training device for the image generation model includes: The first acquisition unit is used to acquire multiple image-text sample pairs, wherein each image-text sample pair includes a background template image, a noise reference image, image description information corresponding to the noise reference image, and a sample object image including a sample object, wherein the noise reference image is obtained by adding noise to a reference image obtained by replacing the reference object in the background template image with the sample object; The first determining unit is configured to determine the features of the sample stitched image based on the noise reference image, the background template image, and the template mask image, wherein the template mask image is obtained by masking the reference object in the background template image; The second determining unit is used to determine the denoising network control information of the image generation model based on the contour features of the reference object in the background template image, the image description information, and the sample object image; The prediction unit is used to perform noise prediction based on the features of the sample stitched image and the control information of the denoising network, through the image generation model, to obtain the noise prediction result corresponding to the noise reference image; The training unit is used to train the image generation model based on the comparison results between the noise reference image and its corresponding noise prediction result in multiple image-text sample pairs.

17. An image generation apparatus, wherein, The image generation device includes: The second acquisition unit is used to acquire an object image including a first object, a background image, and descriptive information, wherein the background image includes a second object, and the descriptive information is used to describe the replacement from the second object to the first object; The third determining unit is used to determine the features of the spliced ​​image based on the background image, the preset noise image, and the mask image, wherein the mask image is obtained by masking the second object in the background image; The fourth determining unit is used to determine the denoising control information of the image generation model based on the contour features of the second object in the background image, the description information, and the object image, wherein the image generation model is generated according to the training method of the image generation model according to any one of claims 1 to 13; An image generation unit is configured to generate an image based on the stitched image features and the denoising control information, using the image generation model, to obtain a target image, wherein the target image is used to indicate the result of replacing the second object in the background image with the first object in the target image.

18. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein, When the processor executes the computer program, it implements the training method of the image generation model according to any one of claims 1 to 13 or the image generation method according to any one of claims 14 to 15.

19. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by the processor, it implements the training method of the image generation model according to any one of claims 1 to 13 or the image generation method according to any one of claims 14 to 15.

20. A computer program product comprising a computer program that is read and executed by a processor of an electronic device, causing the electronic device to perform a training method for an image generation model according to any one of claims 1 to 13 or an image generation method according to any one of claims 14 to 15.