Image synthesis method and device, equipment and storage medium

Through the iterative denoising and redrawing method, the mapping relationship between configuration parameters and noise proportion is used to adjust the feature participation degree, which solves the problem of image synthesis diversity and excessive consumption of computing resources, and achieves efficient and stable image synthesis.

CN120451718APending Publication Date: 2025-08-08TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526307.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The diversity of synthetic images in existing image synthesis methods is poor, and optimization processing will lead to excessive consumption of computing resources or high complexity of the synthesis process, making it difficult to implement and deploy.

Method used

The iterative denoising and redrawing method is used to limit the degree of participation of features in the denoising and redrawing process through the mapping relationship between preset configuration parameters and noise proportion, and adjust the feature participation rate one by one to generate a synthetic image set.

Benefits of technology

It improves the diversity and accuracy of synthetic images, avoids the problems of excessive consumption of computing resources and excessive synthesis process, and enhances the stability and applicability of image synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451718A_ABST
    Figure CN120451718A_ABST
Patent Text Reader

Abstract

The invention provides an image synthesis method and device, equipment and a storage medium, which are used for solving the problem of relatively large synthesis limitation during image synthesis. The method at least comprises the following steps: based on a preset mapping relationship between at least one configuration parameter and a noise ratio, obtaining at least one configuration parameter of this round corresponding to the noise ratio of a previous intermediate feature set obtained by denoising and redrawing of a previous round; each configuration parameter is in negative correlation with the noise ratio and is used for limiting the participation degree of one feature during de-noising and redrawing; according to the at least one current-round configuration parameter, based on the content semantic features, denoising and redrawing are carried out on the previous intermediate feature set, and a current-round intermediate feature set is obtained; and taking the obtained current round of intermediate feature set as the previous intermediate feature set, entering the next round of denoising and redrawing until the final round of denoising and redrawing, and generating the composite image set based on the obtained current round of intermediate feature set. And the synthesis limitation is reduced by adjusting the feature participation degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image synthesis method, apparatus, device and storage medium. Background Art

[0002] With the continuous development of technology, more and more client or server devices can generate synthetic image sets according to text descriptions, so that the synthetic image sets can be applied to perform downstream tasks.

[0003] In related technologies, an image synthesis method may be: using a trained text-based image model to predict a set of synthesized images based on input prompt text.

[0004] However, since the synthetic image set is predicted based on the same prompt text, the diversity of the synthetic images in the synthetic image set is poor and the image details are not rich enough. In order to improve the image diversity of the final synthetic image set, related technologies also perform some optimization processing before and after the text-based graph model processing. However, these optimization processing will bring some limitations to image synthesis, as follows:

[0005] 1. If the prompt text is expanded, such as adding a large number of detailed descriptions or replacing multiple detailed descriptions, the prompt text will have a large amount of text content or text quantity, which will consume a lot of computing resources to process the prompt text, and it is easy to cause crashes such as synthesis failure;

[0006] Second, if after obtaining the set of synthetic images, the image details in each synthetic image are modified,

[0007] This will make the image synthesis process longer, more complex, and difficult to implement and deploy. Summary of the Invention

[0008] The embodiments of the present application provide an image synthesis method, apparatus, device, and storage medium for solving the problem of large synthesis limitations when synthesizing images.

[0009] In a first aspect, an image synthesis method is provided, comprising:

[0010] Obtaining content semantic features of the prompt content, and obtaining an original image feature set including original image features of at least one noise image;

[0011] Based on the content semantic features, the original image feature set is subjected to multiple rounds of denoising and redrawing in an iterative manner to obtain a synthetic image set; wherein each round of denoising and redrawing includes:

[0012] Based on the mapping relationship between at least one preset configuration parameter and the noise ratio, at least one configuration parameter for this round corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing is obtained; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing;

[0013] According to the at least one current round configuration parameter and based on the content semantic features, denoising and redrawing the previous intermediate feature set to obtain the current round intermediate feature set;

[0014] The obtained intermediate feature set of this round is used as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, at which time the synthetic image set is generated based on the obtained intermediate feature set of this round.

[0015] In a second aspect, an image synthesis device is provided, comprising:

[0016] Acquisition module: used for acquiring content semantic features of prompt content and acquiring an original image feature set including original image features of at least one noise image;

[0017] A processing module is configured to perform multiple rounds of denoising and redrawing on the original image feature set in an iterative manner based on the content semantic features to obtain a synthetic image set; wherein each round of denoising and redrawing includes:

[0018] The processing module is specifically configured to: obtain, based on a mapping relationship between at least one preset configuration parameter and the noise ratio, at least one configuration parameter of the current round corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing;

[0019] The processing module is specifically configured to: perform denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the at least one current round configuration parameter to obtain the current round intermediate feature set;

[0020] The processing module is specifically configured to use the obtained intermediate feature set of this round as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, and generate the synthetic image set based on the obtained intermediate feature set of this round.

[0021] Optionally, the configuration parameter includes a configuration weight, and the configuration weight is used to limit the degree of participation of the content semantic feature; and the mapping relationship includes: a weight mapping relationship between the noise ratio and the configuration weight;

[0022] The processing module is specifically used for:

[0023] According to the configuration weights of this round obtained from the weight mapping relationship, based on the content semantic features, the previous intermediate feature set is denoised and redrawn to obtain the intermediate feature set of this round.

[0024] Optionally, the processing module is specifically configured to:

[0025] Denoising the previous intermediate feature set based on the content semantic features according to the configuration weights of this round obtained from the weight mapping relationship to obtain a denoised feature set of this round;

[0026] According to other current-round configuration parameters except the current-round configuration weight among the at least one current-round configuration parameter, the current-round denoising feature set is redrawn to obtain a current-round intermediate feature set.

[0027] Optionally, the configuration parameter includes a configuration threshold, the configuration threshold being used to limit the participation degree of the current round of denoising feature set by limiting the upper limit of similarity between the current round of denoising feature set generated during the denoising and redrawing process and the reference image features of a pre-stored reference image; and the mapping relationship includes: a threshold mapping relationship between the noise ratio and the configuration threshold;

[0028] The processing module is specifically used for:

[0029] Based on the content semantic features, denoising the previous intermediate feature set to obtain a denoising feature set for this round;

[0030] Based on the feature similarity between each current-round denoising feature in the current-round denoising feature set and the reference image feature, combined with the current-round configuration threshold obtained from the threshold mapping relationship, the current-round denoising feature set is redrawn to obtain the current-round intermediate feature set.

[0031] Optionally, the processing module is specifically configured to:

[0032] Obtaining the feature similarity between each denoising feature of this round and the reference image feature in the denoising feature set of this round;

[0033] Redraw the denoising features of this round whose feature similarity is greater than the configuration threshold of this round to obtain the intermediate features of this round;

[0034] The denoising features of this round whose feature similarity is not greater than the configuration threshold of this round are used as the intermediate features of this round, and the obtained intermediate features of this round are arranged to obtain the intermediate feature set of this round.

[0035] Optionally, the configuration parameters include configuration similarity, and the configuration similarity is used to limit the participation degree of the current round of denoising feature set by limiting the lower limit of the similarity between any feature element in the current round of denoising feature set generated during the denoising and redrawing process and any feature element in the reference image features; the mapping relationship includes: a similarity mapping relationship between the noise ratio and the configuration similarity;

[0036] The processing module is specifically used for:

[0037] Based on the content semantic features, denoising the previous intermediate feature set to obtain a denoising feature set for this round;

[0038] Obtaining the element similarity between each feature element in each current round of denoising features and each feature element in the reference image features;

[0039] Based on the obtained similarities of each element and the current round configuration similarity obtained from the similarity mapping relationship, the current round denoising feature set is redrawn to obtain the current round intermediate feature set.

[0040] Optionally, the processing module is specifically configured to:

[0041] For each current-round denoising feature in the current-round denoising feature set, perform the following operations:

[0042] Based on the obtained similarities of each element, determining the number of elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity;

[0043] Redrawing the denoising features of this round based on the number of elements to obtain the intermediate features of this round;

[0044] The intermediate features obtained in this round are arranged to obtain the intermediate feature set of this round.

[0045] Optionally, the processing module is specifically configured to:

[0046] For each feature element in the current round of denoising features, the following operations are performed: based on the element similarity between a feature element in the current round of denoising features and each feature element in the reference image features, a feature element in the reference image features whose element similarity satisfies a similarity condition is used as a target element corresponding to the feature element;

[0047] Based on the element correlation between each feature element in the current round of denoising features and its corresponding target element, the number of feature elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity is determined.

[0048] Optionally, the processing module is specifically configured to:

[0049] For each current-round denoising feature in the current-round denoising feature set, performing the following steps: determining, based on the obtained element similarities, a feature element in the current-round denoising feature whose element similarity is greater than the current-round configuration similarity as a to-be-redrawn element; and determining, in the reference image feature, a feature element whose element similarity to the to-be-redrawn element is greater than the current-round configuration similarity as a reference element;

[0050] Based on the obtained element differences between each to-be-redrawn element and its corresponding reference element, each to-be-redrawn element in the current round of denoising features is adjusted to obtain the current round of intermediate features.

[0051] Optionally, the configuration parameter includes a configuration difference; the configuration difference is used to limit the participation degree of the denoising feature in this round by limiting the difference between the denoising feature in this round before and after redrawing; the mapping relationship includes: a mapping relationship between the noise ratio and the configuration difference;

[0052] The processing module is specifically used for:

[0053] Based on the configuration difference of this round obtained from the difference mapping relationship, each of the elements to be redrawn and its corresponding element difference are fused in the denoising features of this round to obtain the intermediate features of this round.

[0054] Optionally, the acquisition module is specifically configured to:

[0055] When the prompt content is in text form, performing text feature extraction on the prompt content, and using the obtained text semantic features as content semantic features;

[0056] When the prompt content is in the form of an image, performing image feature extraction on the prompt content, and using the obtained image semantic features as content semantic features;

[0057] When the prompt content is in audio form, audio features are extracted from the prompt content, and the obtained audio semantic features are used as content semantic features.

[0058] According to a third aspect, a computer program product is provided, comprising a computer program, which implements the method according to the first aspect when executed by a processor.

[0059] According to a fourth aspect, a computer device is provided, comprising:

[0060] memory for storing computer programs;

[0061] The processor is configured to call the computer program stored in the memory and execute the method according to the first aspect according to the obtained computer program.

[0062] In a fifth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is used to enable a computer to execute the method described in the first aspect.

[0063] In the embodiment of the present application, each configuration parameter is negatively correlated with the noise ratio, and is used to limit the degree of participation of a feature in denoising and redrawing. As the noise ratio becomes smaller and smaller through iteration, the degree of participation of at least one feature related to the denoising and redrawing process becomes higher and higher, so that when the noise ratio is large, the denoising and redrawing process has greater freedom, thereby improving the diversity of the intermediate feature set obtained after each round of denoising and redrawing; when the noise ratio is small, the intermediate feature set obtained after each round of denoising and redrawing contains more image content. At this time, increasing the degree of participation of at least one feature related to the denoising and redrawing process can make the obtained synthetic image set more in line with the requirements of the prompt content, thereby improving the accuracy of image synthesis.

[0064] The entire image synthesis process does not require the expansion of prompt content, avoiding crashes such as synthesis failure due to the large amount of prompt content data; it also does not require detailed modifications, avoiding the problem of large limitations such as the image synthesis process being too long and difficult to implement and deploy.

[0065] Furthermore, a mapping relationship is preset between at least one configuration parameter and the noise ratio. Then, for different application scenarios, different participation levels can be adjusted for each feature, thereby improving the applicability of the image synthesis process. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1A A schematic diagram of a field of the image synthesis method provided in an embodiment of the present application;

[0067] Figure 1B A schematic diagram of a principle of an image synthesis method in related art;

[0068] Figure 1C An application scenario of the image synthesis method provided in the embodiment of the present application;

[0069] Figure 2 A schematic diagram of a flow chart of an image synthesis method provided in an embodiment of the present application;

[0070] Figure 3A A schematic diagram of a principle of an image synthesis method provided in an embodiment of the present application;

[0071] Figure 3B A schematic diagram of the principle of the image synthesis method provided in the embodiment of the present application Figure 2 ;

[0072] Figure 4AA third schematic diagram of a principle of the image synthesis method provided in an embodiment of the present application;

[0073] Figure 4B A fourth schematic diagram of a principle of the image synthesis method provided in an embodiment of the present application;

[0074] Figure 4C A fifth schematic diagram of a principle of the image synthesis method provided in an embodiment of the present application;

[0075] Figure 4D Schematic diagram 6 of a principle of the image synthesis method provided in an embodiment of the present application;

[0076] Figure 5A Schematic diagram 7 of a principle of the image synthesis method provided in an embodiment of the present application;

[0077] Figure 5B A schematic diagram of the principle of the image synthesis method provided in the embodiment of the present application Figure 8 ;

[0078] Figure 6A A schematic diagram of the principle of the image synthesis method provided in the embodiment of the present application Figure 9 ;

[0079] Figure 6B Schematic diagram 10 of a principle of an image synthesis method provided in an embodiment of the present application;

[0080] Figure 7A 11 is a schematic diagram showing a principle of an image synthesis method provided in an embodiment of the present application;

[0081] Figure 7B A schematic diagram showing a principle of an image synthesis method provided in an embodiment of the present application is shown in FIG12;

[0082] Figure 7C 13. Schematic diagram of a principle of an image synthesis method provided in an embodiment of the present application;

[0083] Figure 8 A first structural diagram of an image synthesis device provided in an embodiment of the present application;

[0084] Figure 9 A schematic diagram of the structure of the image synthesis device provided in the embodiment of the present application Figure 2 . DETAILED DESCRIPTION

[0085] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0086] Some of the terms used in the embodiments of the present application are explained below to facilitate understanding by those skilled in the art.

[0087] (1) Diffusion Models:

[0088] Diffusion models uniquely generate high-quality data by gradually adding noise to a dataset and then learning how to reverse this process. This enables diffusion models to create highly accurate and detailed outputs, ranging from realistic images to coherent text sequences. The core function of diffusion models is to gradually degrade data before reconstructing it to its original form or transforming it into something new.

[0089] (2) NegToMe model:

[0090] The NegToMe model, used for image generation, aims to improve the diversity of generated images and mitigate copyright risks through image-driven adversarial guidance. By directly referencing the visual features of an image, the NegToMe model achieves precise and flexible control over image generation, significantly improving the diversity of generated images, particularly in visual feature processing.

[0091] It should be noted that the embodiments of the present application involve operations to obtain prompt content, reference images and other data. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0092] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0093] The following is a brief introduction to the application fields of the image synthesis method provided in the embodiments of the present application.

[0094] With the continuous advancement of technology, more and more clients and servers can generate synthetic image sets based on text descriptions, allowing them to be used to perform downstream tasks. For example, for downstream image creation tasks, after entering prompt text, the generated synthetic image should be both beautiful and compliant with the text, while also being as diverse as possible. Furthermore, for copyrighted images (such as Iron Man and Spider-Man), the generated synthetic image should contain images with similar functions, but with completely different materials.

[0095] For example, see Figure 1A In (1), in the field of design, according to the creative text "design an advertising image for a new environmentally friendly sports drink, in which an athlete is running in a forest full of green plants and fresh air, holding a drink bottle in his hand, and the drink bottle has a unique environmental protection logo", a set of visual synthetic images (in which the faces are fictional, not real faces) can be generated, so that the creative effect can be intuitively evaluated.

[0096] For another example, please refer to Figure 1A (2) In the field of content creation, we can generate a visual composite image set based on the content text of the novel, "drawing a picture of a magical forest with glowing exotic plants, mysterious huts hidden among the trees, and flying magical elves", so that we can match corresponding illustrations for novels of various styles.

[0097] For another example, please refer to Figure 1A In (3), in the field of teaching, we can follow the text in the teaching courseware "Draw the cell structure, the nucleus is represented by blue, the mitochondria is represented by red, and the cytoplasm is represented by light yellow" to generate a visual synthetic image set to assist teaching.

[0098] For related technologies, please refer to Figure 1B ,The image synthesis method can be as follows: using the trained ,text-image model to predict a set of synthetic images based on the ,input prompt text.

[0099] However, since the synthetic image set is predicted based on the same prompt text, the diversity of the synthetic images in the synthetic image set is poor and the image details are not rich enough. In order to improve the image diversity of the final synthetic image set, related technologies also perform some optimization processing before and after the text-based graph model processing. However, these optimization processing will bring some limitations to image synthesis, as follows:

[0100] 1. If the prompt text is expanded, such as adding a large number of detailed descriptions or replacing multiple detailed descriptions, for example, the expanded prompt text is "Draw several overlapping clouds in the image area in the upper right corner, and the clouds gradually become thinner from right to left. Draw a school in the image area in the lower left corner, and the school walls are brick red..."

[0101] For another example, the expanded prompt text includes "a little girl in a yellow dress, a high ponytail, and big eyes", "a little girl in a ballet costume dancing", "a little girl in a school uniform, carrying a schoolbag, and walking", etc.

[0102] Then, the prompt text will have a large amount of text content or text quantity, which will consume a lot of computing resources to process the prompt text, and it will be easy to cause crashes such as synthesis failure;

[0103] Second, if the image details in each synthesized image are modified after obtaining the synthesized image set, the image synthesis process will be longer and more complex, making it difficult to implement and deploy.

[0104] To address the significant limitations of image synthesis, this application proposes an image synthesis method. After obtaining the content semantic features of the prompt content and an original image feature set containing the original image features of at least one noise image, the original image feature set is subjected to multiple rounds of denoising and redrawing in an iterative manner based on the content semantic features to obtain a synthesized image set. Each round of denoising and redrawing includes:

[0105] Based on the mapping relationship between at least one preset configuration parameter and the noise ratio, at least one current-round configuration parameter corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing is obtained; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing. According to at least one current-round configuration parameter, the previous intermediate feature set is denoised and redrawn based on content semantic features to obtain the current-round intermediate feature set. The obtained current-round intermediate feature set is used as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, at which point a synthetic image set is generated based on the obtained current-round intermediate feature set.

[0106] In the embodiment of the present application, each configuration parameter is negatively correlated with the noise ratio, and is used to limit the degree of participation of a feature in denoising and redrawing. As the noise ratio becomes smaller and smaller through iteration, the degree of participation of at least one feature related to the denoising and redrawing process becomes higher and higher, so that when the noise ratio is large, the denoising and redrawing process has greater freedom, thereby improving the diversity of the intermediate feature set obtained after each round of denoising and redrawing; when the noise ratio is small, the intermediate feature set obtained after each round of denoising and redrawing contains more image content. At this time, increasing the degree of participation of at least one feature related to the denoising and redrawing process can make the obtained synthetic image set more in line with the requirements of the prompt content, thereby improving the accuracy of image synthesis.

[0107] The entire image synthesis process does not require the expansion of prompt content, avoiding crashes such as synthesis failure due to the large amount of prompt content data; it also does not require detailed modifications, avoiding the problem of large limitations such as the image synthesis process being too long and difficult to implement and deploy.

[0108] Furthermore, a mapping relationship is preset between at least one configuration parameter and the noise ratio. Then, for different application scenarios, different participation levels can be adjusted for each feature, thereby improving the applicability of the image synthesis process.

[0109] The following describes the application scenarios of the image synthesis method provided by this application.

[0110] Please refer to Figure 1C (The faces in the figure are fictitious, not real faces.) This is a schematic diagram of an application scenario of the image synthesis method provided by this application. This application scenario includes a server 101 and a client 102. Server 101 and client 102 can communicate with each other. The communication method can be wired communication technology, such as connecting via a network cable or serial port cable, or wireless communication technology, such as Bluetooth or wireless fidelity (WIFI), without limitation.

[0111] The client 102 generally refers to a device that can provide prompt content or present a composite image set, such as a terminal device, a third-party application accessible to the terminal device, or a webpage accessible to the terminal device. The server 101 generally refers to a device that can execute an image synthesis method, such as a terminal device or a server.

[0112] Terminal devices include, but are not limited to, mobile phones, computers, smart medical devices, smart home appliances, in-vehicle terminals, or aircraft. Servers include, but are not limited to, cloud servers, local servers, or associated third-party servers. Both client 101 and server 102 can utilize cloud computing to reduce the use of local computing resources; similarly, cloud storage can be used to reduce the use of local storage resources.

[0113] As an embodiment, the server 101 and the client 102 may be the same device, or may be different devices, or may be different devices that share some modules, etc., without specific limitation.

[0114] In response to an input operation triggered by prompt content, client 102 generates an image synthesis request containing the prompt content. Client 102 sends the image synthesis request to server 101. Server 101 receives the image synthesis request from client 102 and obtains the prompt content. Based on the prompt content, server 101 generates a set of synthesized images that match the prompt content. Server 101 returns the synthesized image set to client 102. After receiving the synthesized image set from server 101, client 102 presents the synthesized image set.

[0115] The following is based on Figure 1C , the image synthesis method provided in the embodiment of this application is specifically introduced. Please refer to Figure 2 , which is a flow chart of the image synthesis method provided in an embodiment of the present application.

[0116] S201 : Acquire content semantic features of prompt content, and acquire an original image feature set including original image features of at least one noise image.

[0117] The prompt content is used to indicate the necessary elements contained in the generated synthetic image set, such as at least one of the elements such as people, animals, objects, environment, color, action, posture, material, emotional color, and style.

[0118] The semantic features of the prompt content can be expressed in the form of vectors, matrices, sequences, or other formats, describing the semantics of the prompt content. These features can be acquired using models such as encoders, from storage units, or sent from other devices. The specific acquisition method is not limited. For example, an encoder (ClipText) can be used to extract the semantic features of the prompt content.

[0119] Noise images can be generated using a preset noise generation strategy, according to the size requirements of the synthetic image, or randomly generated, with no specific restrictions. The noise data in different noise images can be the same or different. The number of synthetic images in a synthetic image set can be the same as or different from the number of noise images. Original image features are used to describe the noise image in the form of vectors, matrices, sequences, etc., to facilitate subsequent calculations.

[0120] Please refer to Figure 3A The original image feature set can be arranged in a multi-dimensional form. For example, the size of the original image feature set can be recorded as B×N×D, where B is the number of original image features, N is the number of matrix rows of the original image features, and D is the number of matrix columns of the original image features.

[0121] As an embodiment, the prompt content can be at least one of multiple forms. Through different forms, the drawing content in the composite image can be prompted from different angles such as characters, vision and hearing, thereby improving the flexibility of image synthesis; multiple forms can prompt the drawing content in the composite image from multiple angles, thereby improving the richness of the generated composite image set.

[0122] For example, when the prompt content is in text form, text feature extraction can be performed on the prompt content, and the obtained text semantic features can be used as content semantic features. For example, a character string can be identified from the prompt content, and a text feature extraction network can be used to extract text semantic features from the character string.

[0123] When the prompt content is in image form, image feature extraction can be performed on the prompt content, and the obtained image semantic features can be used as content semantic features. For example, at least one key element can be identified from the prompt content, and an image feature extraction network can be used to extract image semantic features from the at least one key element. For another example, at least one string can be detected from the prompt content, and a text feature extraction network can be used to extract text semantic features from each string as image semantic features.

[0124] When the prompt content is in audio format, audio features are extracted from the prompt content, and the obtained audio semantic features are used as content semantic features. For example, at least one string is identified from the prompt content, and a text feature extraction network is used to extract text semantic features from each string as audio semantic features. For another example, emotional tone is detected from the prompt content, and a semantic feature extraction network is used to extract audio semantic features from the emotional tone.

[0125] Please refer to Figure 3BWhen the prompt content contains sub-content in text form and sub-content in image form, the above method can be used to obtain the text semantic features of the sub-content in text form and the image semantic features of the sub-content in image form, and splice them into content semantic features.

[0126] When the prompt content includes text sub-content and audio sub-content, the above method can be used to respectively obtain text semantic features of the text sub-content and audio semantic features of the audio sub-content, and splice them into content semantic features.

[0127] When the prompt content includes sub-content in image form and sub-content in audio form, the above method can be used to respectively obtain image semantic features of the sub-content in image form and audio semantic features of the sub-content in audio form, and splice them into content semantic features.

[0128] When the prompt content includes sub-content in text form, sub-content in image form, and sub-content in audio form, the above method can be used to obtain the text semantic features of the sub-content in text form, the image semantic features of the sub-content in image form, and the audio semantic features of the sub-content in audio form, and obtain the content semantic features through weighted summation.

[0129] S202 , based on the content semantic features, an iterative method is used to perform multiple rounds of denoising and redrawing on the original image feature set to obtain a synthetic image set.

[0130] After obtaining the content semantic features and the original image feature set, the original image feature set can be subjected to multiple rounds of denoising and redrawing in an iterative manner based on the content semantic features to obtain a synthetic image set. In the first round of denoising and redrawing, each original image feature in the original image feature set is denoised and redrawn based on the content semantic features to obtain the corresponding intermediate features of this round, thereby obtaining the intermediate feature set of this round. The noise proportion (such as the signal-to-noise ratio) in each intermediate feature of this round in the intermediate feature set of this round is determined. When the average value of the noise proportion of each intermediate feature of this round in the intermediate feature set of this round is lower than the denoising target, a synthetic image is generated based on each intermediate feature of this round in the intermediate feature set of this round to obtain a synthetic image set; otherwise, the intermediate feature set of this round is used as the previous intermediate feature set to enter the next round of denoising and redrawing.

[0131] When generating a synthetic image based on the intermediate features of this round, a prediction model such as an image decoder can be used to generate a synthetic image that meets the prompt content based on the intermediate features of this round in the final round.

[0132] Multiple rounds of denoising and redrawing can be implemented using a trained diffusion model, etc., and there is no restriction here.

[0133] The following is an introduction to one round of denoising and redrawing process as an example. Please refer to S2021 to S2023. The other rounds of denoising and redrawing processes are similar and will not be repeated here.

[0134] S2021: Based on a mapping relationship between at least one preset configuration parameter and the noise proportion, obtain at least one configuration parameter of this round corresponding to the noise proportion of the previous intermediate feature set obtained by the previous round of denoising and redrawing.

[0135] Since each original image feature in the original image feature set is a synthetic image set obtained by performing multiple rounds of denoising and redrawing processes based on the same content semantic feature, in order to increase the image difference between the individual synthetic images in the synthetic image set and improve the synthetic diversity, in an embodiment of the present application, a mapping relationship between at least one configuration parameter and the noise proportion is pre-set, and each configuration parameter is negatively correlated with the noise proportion. Each configuration parameter is used to limit the degree of participation of a feature in denoising and redrawing.

[0136] At least one feature is a feature that participates in the calculation in each round of denoising and redrawing, and there is no specific restriction.

[0137] Then, when the noise proportion is greater, the degree of participation of at least one feature in the denoising and redrawing is smaller, which makes the degree of freedom of the denoising and redrawing process higher, and can achieve the purpose of increasing the feature difference between each round of intermediate features in the intermediate feature set obtained in each round, thereby increasing the image difference between the individual synthetic images in the synthetic image set and improving the synthetic diversity.

[0138] The smaller the noise ratio, the greater the participation of at least one feature in the denoising and redrawing. Then, each intermediate feature in the intermediate feature set obtained in each round can better reflect the image content of the synthesized image, so that the denoising and redrawing process can proceed in a direction that is more consistent with the prompt content, thereby achieving the purpose of improving the synthesis accuracy while ensuring the diversity of synthesis.

[0139] Since there is no need to expand the prompt content during the entire multi-round denoising and redrawing process, it not only ensures the diversity of synthesis, but also avoids the synthesis failure and other crashes caused by the excessive amount of prompt content data when each round of denoising and redrawing is based on the prompt content as the standard, thereby improving the stability of image synthesis.

[0140] At the same time, there is no need to modify the details of each synthesized image in turn after outputting the synthesized image set, which avoids the lengthy synthesis process and facilitates the flexible use of the image synthesis method.

[0141] There are many configuration parameters. The following describes four configuration parameters as examples. Other configuration parameters are not listed here one by one.

[0142] Configuration parameter 1: Configuration weight

[0143] Configuration weights are used to limit the involvement of content semantic features. Therefore, the mapping relationship can include a weight mapping relationship between noise ratio and configuration weights. For example, if the content semantic features describe multiple key elements, then when the noise ratio is high, the previous intermediate feature set contains more noise and less image content relevant to the synthesized image set. Therefore, the previous intermediate feature set can be denoised and redrawn based only on the partial features of the content semantic features that describe a key element. This reduces the computational complexity of the current denoising and redrawing process, improves synthesis stability, and increases the image differences between the intermediate features in the current intermediate feature set, thereby improving synthesis diversity.

[0144] When the noise ratio is small, since the previous intermediate feature set contains less noise and has more image content related to the synthesized image set, the previous intermediate feature set can be denoised and redrawn based on all the features of all the key elements described in the content semantic features to ensure that the synthesized image is consistent with the prompt content and improve the synthesis accuracy.

[0145] The weight mapping relationship may indicate the configuration weights corresponding to each of a plurality of noise proportion intervals, or the configuration weights corresponding to each of a plurality of noise proportions, or the functional relationship between the noise proportion and the configuration weight, etc., without specific limitation.

[0146] Please refer to Figure 4A (The image shows a fictitious face, not a real one), the prompt "16 years old, male, black curly hair, wearing round glasses, wearing Gryffindor robes, with a lightning-shaped scar on the face" contains multiple key elements, such as "16 years old", "male", "black curly hair", "round glasses", "Gryffindor robes" and "with a lightning-shaped scar on the face".

[0147] Then, when the noise proportion of the previous intermediate feature set is relatively large, the previous intermediate feature can be denoised and redrawn based only on the partial features corresponding to a key element in the content semantic features (one previous intermediate feature is taken as an example in the figure, and other previous intermediate features are similar) to obtain the intermediate features of this round.

[0148] When the noise ratio of the previous intermediate feature set is relatively small, the previous intermediate feature can be denoised and redrawn based only on the partial features corresponding to multiple key elements in the content semantic features (one previous intermediate feature is taken as an example in the figure, and the other previous intermediate features are similar) to obtain the intermediate features of this round.

[0149] Configuration parameter 2: Configuration threshold

[0150] Configuring a threshold limits the participation of the denoising feature set generated during the denoising redraw process by capping the similarity between the current denoising feature set and the reference image features of the stored reference image. Only the denoising features generated during the denoising redraw process that meet the similarity cap are redrawn, eliminating the need to redraw all denoising features. This reduces computational effort and improves image synthesis stability.

[0151] The mapping relationship can include a threshold mapping relationship between noise percentage and a configured threshold. When the noise percentage is high, each denoising feature in the denoising feature set generated during the denoising and redrawing process contains more noise and less image content related to the synthesized image set. The denoising features are relatively similar to each other and relatively dissimilar to the reference image features. In this case, the upper limit of similarity is small. Since redrawing is only performed when the upper limit of similarity is reached, it avoids misjudgments such as denoising features that need to be redrawn being misjudged as not needing to be redrawn. This improves synthesis accuracy while ensuring synthesis stability.

[0152] When the noise ratio is small, each denoising feature in the denoising feature set generated during the denoising redrawing process contains less noise and more image content related to the synthetic image set. The denoising features of each round are relatively dissimilar. At this time, the upper limit of the similarity is large. At this time, by configuring a threshold to limit the upper limit of the similarity, it is avoided that the redrawing of a large number of denoising features of the current round causes a large amount of calculation, which leads to low synthetic stability.

[0153] The threshold mapping relationship may indicate the configuration thresholds corresponding to multiple noise proportion intervals, or the configuration thresholds corresponding to multiple noise proportions, or the functional relationship between the noise proportion and the configuration threshold, etc., without specific limitation.

[0154] The reference image can be obtained from a storage unit or received from other devices, etc., then the reference image features of the reference image are extracted using a pre-stored image feature extraction network, or extracted after adding noise to the reference image according to the noise amount of the denoising features generated in each round of denoising and redrawing, etc., and there is no specific limitation.

[0155] Alternatively, the reference image may not be obtained, and any denoising feature in the denoising feature set generated in each round of denoising and redrawing may be directly used as the reference image feature of the reference image to increase the difference between the synthetic images in the synthetic image set and improve the synthetic diversity.

[0156] Please refer to Figure 4BTaking the current round denoising feature set containing two current round denoising features (the figure shows a visualized image of the current round denoising features as an example), when the noise proportion of the previous intermediate feature set is large, the noise proportion of the current round denoising feature set obtained after the current round denoising is still relatively large, then each current round denoising feature has a large difference from the reference image, and the similarity is low. At this time, by configuring the threshold to limit the upper limit of the similarity to a smaller value, the synthesis accuracy is improved while ensuring the synthesis stability.

[0157] When the noise ratio of the previous intermediate feature set is relatively small, the noise ratio of the denoised feature set obtained after this round of denoising is relatively small. Then, each denoised feature of this round has a small difference with the reference image, and the similarity is high. At this time, by configuring the threshold to limit the upper limit of the similarity, the stability of image synthesis is improved.

[0158] Configuration parameter three: configuration similarity

[0159] Configuring similarity limits the participation of the denoising feature set in this round by limiting the similarity between any feature element in the denoising feature set generated during the denoising redrawing process and any feature element in the reference image. Feature elements describe the characteristics of an image region. Only feature elements in each denoising feature set that meet the similarity threshold are redrawn, eliminating the need to redraw all feature elements. This reduces computational effort and improves image synthesis stability.

[0160] The mapping relationship can include a similarity mapping relationship between noise ratio and configuration similarity. When the noise ratio is high, each denoising feature in the denoising feature set generated during the denoising and redrawing process contains more noise and less image content related to the synthesized image set. The differences between the feature elements in the denoising feature set are small. In this case, the lower limit of similarity is small. Since only feature elements that meet the lower limit of similarity are redrawn, this avoids misjudgments such as feature elements that need to be redrawn being misjudged as not needing to be redrawn. This improves synthesis accuracy while ensuring synthesis stability.

[0161] When the noise ratio is small, each denoising feature in the denoising feature set generated during the denoising redrawing process contains less noise and more image content related to the synthetic image set. The differences between the feature elements in the denoising feature are relatively large. At this time, by configuring the similarity limit, the lower limit of the similarity is larger to avoid the large amount of calculation caused by redrawing a large number of feature elements, which leads to lower synthetic stability.

[0162] The similarity mapping relationship may indicate the configuration similarities corresponding to multiple noise proportion intervals, or the configuration similarities corresponding to multiple noise proportions, or the functional relationship between noise proportion and configuration similarity, etc., without specific limitation.

[0163] Please refer to Figure 4C Taking the current round denoising feature set containing two current round denoising features (the figure shows the visual image of the current round denoising features as an example, and the feature elements are shown in thick solid rectangular boxes) as an example, when the noise proportion of the previous intermediate feature set is large, the noise proportion of the current round denoising feature set obtained after the current round denoising is still relatively large, then each current round denoising feature has a large difference from the reference image, and the similarity is low. At this time, by configuring the threshold to limit the upper limit of the similarity to a smaller value, the synthesis accuracy is improved while ensuring the synthesis stability.

[0164] When the noise ratio of the previous intermediate feature set is relatively small, the noise ratio of the denoised feature set obtained after this round of denoising is relatively small. Then, each denoised feature of this round has a small difference with the reference image, and the similarity is high. At this time, by configuring the threshold to limit the upper limit of the similarity, the stability of image synthesis is improved.

[0165] Configuration parameter 4: Configuration difference

[0166] The configuration difference is used to limit the participation of the denoising features in this round by limiting the difference between the features before and after redrawing.

[0167] The mapping relationship includes: the difference between noise proportion and configuration difference. Each denoising feature in the current round is redrawn according to the configuration difference. When the noise proportion is large, since the denoising features of each round are relatively similar, the denoising features of this round can be slightly extended to reduce the computational cost and improve the synthesis stability. When the noise proportion is small, since the denoising features of this round can more accurately reflect the image content related to the synthesized image, the denoising features of this round can be significantly extended to improve the synthesis diversity.

[0168] The difference mapping relationship may indicate the configuration differences corresponding to multiple noise proportion intervals, or the configuration differences corresponding to multiple noise proportions, or the functional relationship between the noise proportion and the configuration differences, etc., without specific limitation.

[0169] Based on the mapping relationship, at least one of the above four configuration parameters can be obtained, and other configuration parameters can also be obtained, again without limitation.

[0170] Please refer to Figure 4DTaking a denoising feature of this round (the figure shows a visualized image of the denoising feature of this round) as an example, when the noise proportion of the previous intermediate feature set is large, the noise proportion of the denoising feature set of this round obtained after this round of denoising is still relatively large, then there is a large difference between the denoising feature of this round and the reference image. At this time, the configuration difference is small, so the redrawing amplitude of the denoising feature of this round is small, and the intermediate feature of this round can still maintain a large difference with the reference image, thereby improving the synthesis stability while ensuring the synthesis accuracy.

[0171] When the noise proportion of the previous intermediate feature set is relatively small, the noise proportion of the denoised feature set obtained after this round of denoising is relatively small. Then each denoised feature of this round has a small difference from the reference image and a high similarity. At this time, the configuration difference is large, which makes the redrawing amplitude of the denoised feature of this round larger. The intermediate features obtained in this round can tend to be more different from the reference image, thereby improving the synthesis diversity.

[0172] S2022: According to at least one configuration parameter of this round and based on the semantic features of the content, denoise and redraw the previous intermediate feature set to obtain the intermediate feature set of this round.

[0173] After obtaining at least one configuration parameter for this round, the previous intermediate feature set can be denoised and redrawn based on the content semantic features according to the at least one configuration parameter for this round to obtain the intermediate feature set for this round. This reduces the computational effort during each denoising and redrawing round, improves synthesis stability, and increases the differences between synthesized images, improving synthesis diversity.

[0174] As an example, during each round of denoising and redrawing, the attention module can first perform feature transformation on each previous intermediate feature in the previous intermediate feature set, extracting more abstract semantic features from the previous intermediate features to obtain the current round's transformed features. Then, based on at least one current round configuration parameter and content semantic features, the transformed feature set is denoised and redrawn to obtain the current round's intermediate feature set. The attention module can include self-attention layers and cross-attention layers, etc., without specific limitation.

[0175] The following describes the denoising and redrawing process based on the four configuration parameters introduced above. The denoising and redrawing processes based on other configuration parameters are not listed here one by one.

[0176] Denoising and redrawing process 1:

[0177] According to the configuration weights of this round obtained from the weight mapping relationship, the previous intermediate feature set is denoised and redrawn based on the content semantic features to obtain the intermediate feature set of this round.

[0178] Because this round of configuration weights limits the involvement of content semantic features in the denoising and redrawing process, when the noise ratio is high, content semantic features are used sparingly, reducing computational effort and improving synthesis stability. When the noise ratio is low, content semantic features are used heavily, ensuring that the synthesized image matches the prompt content and ensuring synthesis accuracy.

[0179] The denoising and redrawing process may include: predicting noise data for each previous intermediate feature in the previous intermediate feature set based on content semantic features according to the current round configuration weights; removing the corresponding noise data from each previous intermediate feature in the previous intermediate feature set to obtain a corresponding denoised feature for the current round, thereby obtaining a denoised feature set for the current round; and performing feature transformation on each denoised feature for the current round using a multi-layer perceptron to obtain a corresponding intermediate feature for generating a composite image. The denoising and redrawing process may also be other processes, which are not limited here.

[0180] Please refer to Figure 5A Taking the configuration weight (Classifier-Free Guidance, cfg) as an example, the weight mapping relationship can be expressed as follows: when the noise ratio is not greater than the min_cfg_end threshold, the configuration weight is 5, and the content semantic feature participation is small; when the noise ratio is greater than the min_cfg_end threshold, the configuration weight is 1, and the content semantic feature participation is large.

[0181] Then, when the noise ratio is greater than the min_cfg_end threshold, the noise data can be predicted based on only one key element "male" in the prompt content, and the noise data can be removed from the previous intermediate feature to obtain the denoising feature of this round.

[0182] cfg primarily represents the degree of compliance with the prompt content. Higher cfg values indicate greater compliance. Using a segmented cfg reduces the influence of prompt content in the early stages of denoising, achieving more diverse initial compositions. This diversity is more pronounced. Lowering cfg in the early stages of denoising reduces the likelihood of composite image corruption. Lowering cfg only takes effect in the early stages of denoising, ensuring compliance with prompt content.

[0183] As an example, after denoising the previous intermediate feature set based on content semantic features according to the current round configuration weights to obtain the current round denoised feature set, if other configuration parameters are obtained, the current round denoised feature set can be redrawn according to at least one current round configuration parameter other than the current round configuration weights to obtain the current round intermediate feature set. This redrawing process can push similar current round denoising features further apart, increase the differences between the individual synthetic images in the generated synthetic image set, and improve synthetic diversity.

[0184] Denoising and redrawing process 2:

[0185] Based on the content semantic features, the previous intermediate feature set is denoised to obtain the denoised feature set of this round (this step can also refer to the corresponding steps in the denoising and redrawing process one, thereby realizing the combination of the denoising and redrawing processes one and two, etc., which is not limited here).

[0186] Based on the feature similarity between each denoising feature in this round and the reference image feature in this round of denoising feature set, combined with the configuration threshold of this round obtained from the threshold mapping relationship, the denoising feature set of this round is redrawn to obtain the intermediate feature set of this round.

[0187] Before determining the feature similarity, each denoising feature in the denoising feature set of this round can be normalized to make the denoising features of each round comparable. The feature similarity between each denoising feature in the denoising feature set of this round and the reference image feature can be determined. The feature similarity can be determined from a global perspective, such as using Euclidean distance or other methods; it can also be determined from a local perspective, such as determining the similarity between each two local features, thereby determining the feature similarity between the denoising feature of this round and the reference image feature (a method for determining feature similarity will be introduced as an example later, and other methods for determining feature similarity from a local perspective are not listed here one by one), etc., without specific restrictions. Therefore, the denoising feature set of this round can be redrawn in combination with the configuration threshold of this round to obtain the intermediate feature set of this round.

[0188] Since the purpose of redrawing is to increase the difference between the intermediate features of each current round in the obtained intermediate feature set, that is, to push away, therefore, through the feature similarity between each denoised feature of this round and the reference image feature in the current round denoising feature set, combined with the current round configuration threshold obtained from the threshold mapping relationship, the similarity between the denoised features of each current round can be judged with a quantitative value, thereby achieving targeted pushing away and improving the synthesis diversity while ensuring the synthesis stability.

[0189] As an embodiment, targeted pushing out can be to first obtain the feature similarity between each denoising feature of this round and the reference image feature in the denoising feature set of this round. If the feature similarity is greater than the configuration threshold of this round, it means that the denoising feature of this round is very similar to the reference image feature. Then, in order to ensure synthetic diversity, the denoising features of this round whose feature similarity is greater than the configuration threshold of this round can be redrawn to achieve feature pushing out and obtain the intermediate features of this round.

[0190] If the feature similarity is not greater than the configuration threshold of this round, it means that the denoising features of this round are not very similar to the reference image features. In order to improve the synthesis stability, the denoising features of this round with feature similarity not greater than the configuration threshold of this round can be not redrawn, and the denoising features of this round with feature similarity not greater than the configuration threshold of this round can be directly used as the intermediate features of this round.

[0191] Therefore, the intermediate feature set of this round can be obtained by arranging the intermediate features of each round.

[0192] For the same prompt content, the generated composite image will always have some similar visual features, such as the face area, the solid color area in the background, etc. If you push all of them away without distinguishing them, it will make the composite image process more prone to collapse. Therefore, please refer to Figure 5B By determining the feature similarity between the denoising features of this round (shown as a visual image in the figure) and the reference image features (shown as a reference image in the figure), it is judged whether the feature similarity is not greater than the configuration threshold sim_th of this round. If it is not greater than sim_th, it means that the similarity between the denoising features of this round and the reference image features is low. If redrawing is performed here, it is easy to cause the collapse of the image synthesis process. Therefore, redrawing can be omitted and the denoising features of this round can be used as the intermediate features of this round; otherwise, redrawing is performed to obtain the original intermediate features to ensure synthesis diversity.

[0193] Denoising and redrawing process three:

[0194] Based on the content semantic features, the previous intermediate feature set is denoised to obtain the current round denoised feature set (this step can also refer to the corresponding steps in the denoising and redrawing process one, thereby realizing the combination of denoising and redrawing processes one and three, etc., which is not limited here). The element similarity between each feature element in each current round denoising feature and each feature element in the reference image feature is obtained. Based on the obtained element similarities and the current round configuration similarity obtained from the similarity mapping relationship, the current round denoising feature set is redrawn to obtain the current round intermediate feature set.

[0195] Through the element similarity between feature elements, the similarity between each image region in the upcoming composite image and each image region in the reference image can be judged. From a local perspective, the diversity between the upcoming composite image and the reference image can be measured, which can effectively improve the synthesis diversity.

[0196] As an embodiment, for each denoising feature in the current round of denoising feature set, it is also possible to determine whether to redraw the denoising feature in the current round based on element similarity. For example, based on the obtained element similarities, the number of feature elements in the denoising feature in the current round whose element similarity is greater than the configuration similarity of the current round is determined. The larger the number of elements, the more similar the denoising feature in the current round is to the reference image feature, and the more redrawing is needed; the smaller the number of elements, the less similar the denoising feature in the current round is to the reference image feature, and the less redrawing is needed. Therefore, the denoising feature in the current round can be redrawn based on the number of elements to obtain the intermediate features of the current round. In this way, the intermediate features obtained in the current round can be arranged to obtain the intermediate feature set of the current round.

[0197] As an embodiment, for each denoising feature in the current round of denoising features, the feature similarity between the current round of denoising features and the reference image features can be determined not only based on operations such as Euclidean distance, but also based on the element similarity between each feature element in the current round of denoising features and each feature element in the reference image features.

[0198] For example, based on the number of feature elements in the current round of denoising features whose element similarity with the feature elements representing the same image region in the reference image features is greater than the configuration similarity of the current round, the feature similarity between the current round of denoising features and the reference image features is determined, and the number of elements is positively correlated with the feature similarity.

[0199] Please refer to Figure 6A , the feature element is denoted as token. When the noise level in the early stage of denoising and redrawing is relatively high, the similarity between each token in the current round of denoising features is relatively high. As the multi-round denoising and redrawing progresses, the features related to the image content of the synthesized image in the current round of denoising features become clearer, and the similarity between each token in the current round of denoising features in the later stage of denoising and redrawing is relatively low. However, if the fixed configuration similarity is set too large, the diversity will not be obvious; if it is set too small, the process of synthesizing the image is likely to crash. Therefore, a variable configuration similarity can be adopted. The noise ratio can be divided into multiple noise ratio intervals, denoted as [t start , t end_0 , [t end_0 , t end_1 …, and then when the noise ratio t is in different noise ratio intervals, the corresponding configuration similarity is used. Each configuration similarity is denoted as merge_th0, merge_th1, …. Among them, merge_th0 < merge_th1 < merge_th2, …, thus, it can adapt to the noise levels of different noise ratios, not only enhancing the diversity but also reducing the probability of the process of synthesizing the image crashing.

[0200] As an embodiment, when determining the number of elements, in addition to based on the corresponding image region, it can also be determined based on similarity conditions. The similarity conditions can be an element similarity interval, or the maximum value of each element similarity, or the element similarity corresponding to the nth ranking in each element similarity, etc., without specific limitation.

[0201] For each feature element in the current round of denoising features, the following operations are performed: Based on the element similarity between a feature element in the current round of denoising features and each feature element in the reference image features, the feature element in the reference image features that satisfies the similarity condition is determined as the target element corresponding to that feature element. For example, based on the element similarity between a feature element in the current round of denoising features and each feature element in the reference image features, the feature element in the reference image features that corresponds to the maximum element similarity is determined as the target element corresponding to that feature element.

[0202] Therefore, based on the element correlation between each feature element in the current round of denoising features and its corresponding target element, the number of feature elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity can be determined.

[0203] When the image regions are not limited, it is possible to avoid the situation where each synthetic image in the synthetic image set is only a plurality of permutations and combinations of multiple image regions, thereby effectively improving the synthetic diversity.

[0204] As an embodiment, when redrawing the denoising feature set of this round based on the obtained similarity of each element and the similarity of the configuration of this round, taking into account the inevitable existence of similar image areas in multiple synthetic images, such as background, skin, color tone, etc., it is possible to redraw only the feature elements in each denoising feature of this round whose element similarity reaches the similarity of the configuration of this round, so as to avoid the low synthetic stability caused by large-scale redrawing.

[0205] For each denoising feature in the current round of denoising, the following steps are performed: based on the obtained element similarities, determine that the feature elements in the current round of denoising whose element similarity is greater than the current round of configuration similarity are the elements to be redrawn; and determine that the feature elements in the reference image whose element similarity to the element to be redrawn is greater than the current round of configuration similarity are the reference elements. Thus, based on the obtained element differences between each element to be redrawn and its corresponding reference element, each element to be redrawn in the current round of denoising can be adjusted to obtain the intermediate features of the current round.

[0206] In conjunction with the previous description, if the target element corresponding to each feature element in the current round of denoising features is obtained in the reference image features, then the reference element may or may not be the target element. Alternatively, based on the element similarity between each feature element and its corresponding target element, the feature element in the current round of denoising features whose element similarity is greater than the current round of configuration similarity is determined to be the element to be redrawn. Therefore, based on the obtained element differences between each element to be redrawn and its corresponding target element, each element to be redrawn in the current round of denoising features can be adjusted separately to obtain the intermediate features of the current round.

[0207] Denoising and redrawing process four:

[0208] Based on the configuration difference of this round obtained from the difference mapping relationship, each element to be redrawn and its corresponding element difference are fused in the denoising features of this round to obtain the intermediate features of this round.

[0209] When adjusting the elements to be redrawn in the current round of denoising features, a fixed difference can be used as a fusion weight. Each element to be redrawn and its corresponding element difference can be fused together in the current round of denoising features to obtain the current round of intermediate features. To further improve synthesis stability, the current round of configuration difference can also be obtained from the difference mapping relationship. This configuration difference can be used as a fusion weight. Each element to be redrawn and its corresponding element difference can be fused together in the current round of denoising features to obtain the current round of intermediate features.

[0210] Please refer to Figure 6B Determine the elemental similarity between each feature element in the current denoising feature (shown as a circle with a diagonal background in the figure) and each feature element in the reference image feature (shown as a circle with a checkered background in the figure). For each feature element in the current denoising feature, select the feature element with the greatest elemental similarity from the reference image features as the target element, and establish an association between the feature element, the target element, and the elemental similarity.

[0211] The number of associations is counted as the number of elements. The ratio of the number of elements to the total number of feature elements in the current round of denoising features is determined as the feature similarity between the current round of denoising features and the reference image features. When the feature similarity is greater than sim_th, the current round configuration similarity corresponding to the noise ratio is determined from the similarity mapping relationship, merge_thi. Feature elements in the current round of denoising features with an element similarity greater than merge_thi in the association relationship are selected as the elements to be redrawn, and the element difference between the redrawn elements and the corresponding target elements is determined.

[0212] Please refer to formula (1). In this round of denoising features, based on the configuration difference α t , merge each element to be redrawn O src The difference between its corresponding element (O src -O target ), and obtain the intermediate features of this round.

[0213] O merge =O src +α t (O src -O target ) (1)

[0214] S2023: Use the obtained intermediate feature set of this round as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, and generate a synthetic image set based on the obtained intermediate feature set of this round.

[0215] After obtaining the intermediate feature set of this round, the noise ratio of the intermediate feature set of this round can be determined. If the noise ratio reaches the denoising target, then this round can be used as the final round of denoising redrawing, and a synthetic image set is generated based on the intermediate feature set obtained in this round.

[0216] If the noise ratio does not reach the denoising target, the intermediate feature set obtained in this round can be used as the previous intermediate feature set, and the noise ratio of the intermediate feature set in this round can be set as the noise ratio of the previous intermediate feature set to enter the next round of denoising and redrawing.

[0217] After obtaining the intermediate feature set of this round, it can also be determined whether the current round reaches the target round. If it reaches the target round, then this round can be used as the final round of denoising and redrawing, and a synthetic image set is generated based on the obtained intermediate feature set of this round.

[0218] If the target round is not reached, the intermediate feature set obtained in this round can be used as the previous intermediate feature set, and the noise ratio of the previous intermediate feature set is determined to enter the next round of denoising and redrawing.

[0219] As an embodiment, the mapping relationship can be configured in a hyperparameter, with the noise ratio of the original image features of the noise image being recorded as t start , as the starting time step of multiple rounds of denoising and redrawing, the noise ratio of the generated synthetic image set is recorded as t end , as the end time step of multiple rounds of denoising and redrawing, for example, t start =1,t end =0. Then, the similarity of each configuration is recorded as [t start , t end_0 ],[t end_0 , t end_1 ]…the corresponding configuration similarities are merge_th0, merge_th1,…; the configuration threshold is recorded as sim_th; the noise ratio threshold of the two configuration weights is recorded as min_cfg_end.

[0220] Please refer to the following Figure 7A , which is a schematic diagram of an image synthesis structure, includes a text encoder, a trained diffusion module (such as the NegToMe model), and an image decoder. The prompt content is input into the text encoder to extract content semantic features, which are then input into the diffusion module for denoising and redrawing. During the denoising and redrawing process, any denoising feature from the current round of denoising feature set can be used as the reference image feature, or reference image features obtained by processing the reference image using other models, etc., without specific restrictions. The intermediate feature set obtained in the final round is input into the image decoder to generate a synthetic image set.

[0221] Taking the prompt content of "16 years old, male, black curly hair, wearing round glasses, wearing Gryffindor robes, and a lightning-shaped scar on the face" as an example, the synthetic image set obtained based on at least one configuration parameter that changes with the change of noise proportion provided in the embodiment of the present application is compared with the synthetic image set obtained by using fixed parameters. Please refer to Figure 7B (The face in the picture is a fictitious one, not a real one).

[0222] It can be seen that in the set of synthetic images (including synthetic image A, synthetic image B, synthetic image C and synthetic image D) obtained based on at least one variable configuration parameter, there are significant differences between the synthetic images in terms of light and shadow, character face shape, color tone, background, clothing details, skin, shooting angle, hairstyle, etc., which makes the synthetic image set diverse.

[0223] However, the synthetic image set obtained based on fixed parameters (including synthetic image E, synthetic image F, synthetic image G and synthetic image H) has poor diversity in terms of similar backgrounds, clothing, facial shapes, hairstyles, lighting, tones, and shooting angles. In synthetic image E, there is even a case of synthesis collapse.

[0224] The following takes the prompt content "25 years old, female, hot pot restaurant owner, wearing a casual apron" as an example to compare the synthetic image set obtained based on at least one configuration parameter that changes with the noise ratio provided in the embodiment of the present application, the synthetic image set obtained by using fixed parameters, and the synthetic image set obtained without involving at least one configuration parameter provided in the embodiment of the present application. Please refer to Figure 7C (The face in the picture is a fictional one, not a real one).

[0225] It can be seen that in the set of synthetic images (including synthetic image A, synthetic image B, synthetic image C and synthetic image D) obtained based on at least one variable configuration parameter, there are significant differences between the synthetic images in terms of light and shadow, color tone, background, clothing style, skin, apron material, etc., which makes the synthetic image set diverse.

[0226] However, the set of synthetic images obtained based on fixed parameters (including synthetic image E, synthetic image F, synthetic image G and synthetic image H) has limited diversity among the synthetic images. For example, there are differences in composition and background, while many aspects such as color tone, clothing style, skin, and apron material are still relatively similar; and it is easy to collapse, generating noisy synthetic images. For example, synthetic images E and synthetic images H show synthetic collapse.

[0227] The set of synthetic images obtained when at least one configuration parameter provided in the embodiments of the present application is not involved (including synthetic image I, synthetic image J, synthetic image K and synthetic image L) has poor diversity among the synthetic images. For example, the color tone, composition, background, clothing style, skin, apron material, apron wearing method, etc. are relatively similar, and the synthetic effect is relatively single.

[0228] Based on the same inventive concept, the embodiment of the present application provides an image synthesis device that can realize the functions corresponding to the aforementioned image synthesis method. Figure 8 , the device includes an acquisition module 801 and a processing module 802, wherein:

[0229] Acquisition module 801: used to acquire content semantic features of prompt content and acquire an original image feature set including original image features of at least one noise image;

[0230] Processing module 802 is used to perform multiple rounds of denoising and redrawing on the original image feature set in an iterative manner based on the content semantic features to obtain a synthetic image set. Each round of denoising and redrawing includes:

[0231] The processing module 802 is specifically configured to: obtain, based on a mapping relationship between at least one preset configuration parameter and the noise ratio, at least one configuration parameter for the current round corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing;

[0232] The processing module 802 is specifically configured to: perform denoising and redrawing on the previous intermediate feature set based on the content semantic features according to at least one configuration parameter of the current round to obtain the intermediate feature set of the current round;

[0233] The processing module 802 is specifically configured to use the intermediate feature set obtained in this round as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, and generate a synthetic image set based on the intermediate feature set obtained in this round.

[0234] In a possible embodiment, the configuration parameters include configuration weights, and the configuration weights are used to limit the degree of participation of content semantic features; and the mapping relationship includes: a weight mapping relationship between noise proportion and the configuration weights;

[0235] The processing module 802 is specifically configured to:

[0236] According to the configuration weights of this round obtained from the weight mapping relationship, the previous intermediate feature set is denoised and redrawn based on the content semantic features to obtain the intermediate feature set of this round.

[0237] In a possible embodiment, the processing module 802 is specifically configured to:

[0238] According to the configuration weights of this round obtained from the weight mapping relationship, based on the content semantic features, the previous intermediate feature set is denoised to obtain the denoised feature set of this round;

[0239] According to at least one configuration parameter of this round, other configuration parameters of this round except the configuration weight of this round, the denoising feature set of this round is redrawn to obtain the intermediate feature set of this round.

[0240] In one possible embodiment, the configuration parameters include a configuration threshold, the configuration threshold being used to limit the participation degree of the current round of denoising feature set by limiting the upper limit of similarity between the current round of denoising feature set generated during the denoising and redrawing process and the reference image features of a pre-stored reference image; and the mapping relationship includes: a threshold mapping relationship between the noise ratio and the configuration threshold;

[0241] The processing module 802 is specifically configured to:

[0242] Based on the content semantic features, the previous intermediate feature set is denoised to obtain the denoised feature set of this round;

[0243] Based on the feature similarity between each denoising feature in this round and the reference image feature in this round of denoising feature set, combined with the configuration threshold of this round obtained from the threshold mapping relationship, the denoising feature set of this round is redrawn to obtain the intermediate feature set of this round.

[0244] In a possible embodiment, the processing module 802 is specifically configured to:

[0245] Obtain the feature similarity between each denoising feature of this round and the reference image feature in the denoising feature set of this round;

[0246] Redraw the denoising features of this round whose feature similarity is greater than the configured threshold of this round to obtain the intermediate features of this round;

[0247] The denoising features of this round whose feature similarity is not greater than the configuration threshold of this round are used as the intermediate features of this round, and the intermediate features obtained in this round are arranged to obtain the intermediate feature set of this round.

[0248] In a possible embodiment, the configuration parameters include configuration similarity, which is used to limit the participation of the current round of denoising feature set by limiting the lower limit of similarity between any feature element in the current round of denoising feature set generated during the denoising and redrawing process and any feature element in the reference image features; the mapping relationship includes: a similarity mapping relationship between the noise ratio and the configuration similarity;

[0249] The processing module 802 is specifically configured to:

[0250] Based on the content semantic features, the previous intermediate feature set is denoised to obtain the denoised feature set of this round;

[0251] Obtain the element similarity between each feature element in each denoising feature of this round and each feature element in the reference image feature;

[0252] Based on the obtained similarity of each element and the configuration similarity of this round obtained from the similarity mapping relationship, the denoising feature set of this round is redrawn to obtain the intermediate feature set of this round.

[0253] In a possible embodiment, the processing module 802 is specifically configured to:

[0254] For each denoising feature in the denoising feature set, perform the following operations:

[0255] Based on the obtained similarities of each element, determine the number of elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity;

[0256] Redraw the denoising features of this round based on the number of elements to obtain the intermediate features of this round;

[0257] The intermediate features obtained in this round are arranged to obtain the intermediate feature set of this round.

[0258] In a possible embodiment, the processing module 802 is specifically configured to:

[0259] For each feature element in the current denoising feature, perform the following operations: based on the element similarity between a feature element in the current denoising feature and each feature element in the reference image feature, take the feature element in the reference image feature that satisfies the similarity condition as the target element corresponding to that feature element;

[0260] Based on the element correlation between each feature element in the current round of denoising features and its corresponding target element, the number of feature elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity is determined.

[0261] In a possible embodiment, the processing module 802 is specifically configured to:

[0262] For each denoising feature in the denoising feature set, the following steps are performed: based on the obtained element similarities, determining as the element to be redrawn the feature element in the denoising feature of the current round whose element similarity is greater than the configuration similarity of the current round; and determining as the reference element the feature element in the reference image whose element similarity to the element to be redrawn is greater than the configuration similarity of the current round;

[0263] Based on the obtained element differences between each to-be-redrawn element and its corresponding reference element, each to-be-redrawn element in the current round of denoising features is adjusted to obtain the current round of intermediate features.

[0264] In a possible embodiment, the configuration parameters include a configuration difference; the configuration difference is used to limit the participation of the denoising features in the current round by limiting the difference between the denoising features before and after redrawing; the mapping relationship includes: a mapping relationship between the noise ratio and the configuration difference;

[0265] The processing module 802 is specifically configured to:

[0266] Based on the configuration difference of this round obtained from the difference mapping relationship, each element to be redrawn and its corresponding element difference are fused in the denoising features of this round to obtain the intermediate features of this round.

[0267] In a possible embodiment, the acquisition module 801 is specifically configured to:

[0268] When the prompt content is in text form, extract text features from the prompt content and use the obtained text semantic features as content semantic features;

[0269] When the prompt content is in the form of an image, image features are extracted from the prompt content, and the obtained image semantic features are used as content semantic features;

[0270] When the prompt content is in audio form, audio features are extracted from the prompt content, and the obtained audio semantic features are used as content semantic features.

[0271] Please refer to Figure 9 , is a computer device 900 provided in an embodiment of the present application, the computer device 900 can be, for example, Figure 1C The client 102 or server 101 in the data storage program. The current version and historical versions of the data storage program and the application software corresponding to the data storage program can be installed on a computer device 900, which includes a processor 980 and a memory 920. In some embodiments, the computer device 900 may include a display unit 940, which includes a display panel 941 for displaying a user interactive operation interface, etc.

[0272] In a possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0273] The processor 980 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 980 reads a data storage program or file, thereby running the data storage program on the computer device 900 and displaying a corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) to perform related operations to implement the technical solutions provided in the embodiments of the present application.

[0274] The memory 920 generally includes internal memory and external memory. The internal memory may be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory may be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 920 is used to store computer programs and other data. The computer program includes an application corresponding to each client, etc. Other data may include data generated after the operating system or application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, the computer program is stored in the memory 920, and the processor 980 executes the computer program in the memory 920 to implement any of the methods discussed in the previous figure.

[0275] The display unit 940 is used to receive input digital information, character information, or contact touch operations / contactless gestures, and to generate signal input related to user settings and function control of the computer device 900. Specifically, in the embodiment of the present application, the display unit 940 may include a display panel 941. The display panel 941, such as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or on the display panel 941) and drive corresponding connected devices according to a pre-set program.

[0276] In one possible embodiment, the display panel 941 may include a touch detection device and a touch controller. The touch detection device detects the player's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 980. The touch controller can also receive and execute commands from the processor 980.

[0277] The display panel 941 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, in some embodiments, the computer device 900 may further include an input unit 930. The input unit 930 may include an image input device 931 and other input devices 932. The other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, and the like.

[0278] In addition to the above, the computer device 900 may also include a power supply 990 for powering other modules, an audio circuit 960, a near-field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an accelerometer, a light sensor, a pressure sensor, etc. The audio circuit 960 specifically includes a speaker 961 and a microphone 962. For example, the computer device 900 can use the microphone 962 to collect the user's voice and perform corresponding operations.

[0279] As an embodiment, the number of the processors 980 may be one or more, and the processor 980 and the memory 920 may be coupled or relatively independently configured.

[0280] As an example, Figure 9 The processor 980 in the embodiment can be used to implement the following Figure 8 The functions of the acquisition module 801 and the processing module 802 in .

[0281] As an example, Figure 9 The processor 980 in can be used to implement the corresponding functions of the server or terminal device discussed above.

[0282] Those skilled in the art will appreciate that all or part of the steps of implementing the above-mentioned method embodiments may be accomplished by a computer program. The aforementioned computer program may be stored in a computer-readable storage medium. When the computer program is executed, it executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0283] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, for example, through a computer program product, which is stored in a storage medium and includes a computer program for enabling a computer device to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0284] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. An image synthesis method, characterized in that: include: Obtaining content semantic features of the prompt content, and obtaining an original image feature set including original image features of at least one noise image; Based on the content semantic features, the original image feature set is subjected to multiple rounds of denoising and redrawing in an iterative manner to obtain a synthetic image set; wherein each round of denoising and redrawing includes: Based on the mapping relationship between at least one preset configuration parameter and the noise ratio, at least one configuration parameter for this round corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing is obtained; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing; According to the at least one current round configuration parameter and based on the content semantic features, denoising and redrawing the previous intermediate feature set to obtain the current round intermediate feature set; The obtained intermediate feature set of this round is used as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, at which time the synthetic image set is generated based on the obtained intermediate feature set of this round.

2. The method according to claim 1, characterized in that The configuration parameters include configuration weights, and the configuration weights are used to limit the degree of participation of the content semantic features; and the mapping relationship includes: a weight mapping relationship between the noise ratio and the configuration weights; The step of performing denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the at least one current round configuration parameter to obtain the current round intermediate feature set includes: According to the configuration weights of this round obtained from the weight mapping relationship, based on the content semantic features, the previous intermediate feature set is denoised and redrawn to obtain the intermediate feature set of this round.

3. The method according to claim 2, characterized in that The step of performing denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the current round configuration weights obtained from the weight mapping relationship to obtain the current round intermediate feature set includes: Denoising the previous intermediate feature set based on the content semantic features according to the configuration weights of this round obtained from the weight mapping relationship to obtain a denoised feature set of this round; According to other current-round configuration parameters except the current-round configuration weight among the at least one current-round configuration parameter, the current-round denoising feature set is redrawn to obtain a current-round intermediate feature set.

4. The method according to claim 1, wherein The configuration parameters include a configuration threshold, the configuration threshold being used to limit the participation of the current round of denoising feature set generated during the denoising and redrawing process by limiting the upper limit of similarity between the current round of denoising feature set and the reference image features of the pre-stored reference image; and the mapping relationship includes: a threshold mapping relationship between the noise ratio and the configuration threshold; The step of performing denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the at least one current round configuration parameter to obtain the current round intermediate feature set includes: Based on the content semantic features, denoising the previous intermediate feature set to obtain a denoising feature set for this round; Based on the feature similarity between each current-round denoising feature in the current-round denoising feature set and the reference image feature, combined with the current-round configuration threshold obtained from the threshold mapping relationship, the current-round denoising feature set is redrawn to obtain the current-round intermediate feature set.

5. The method according to claim 4, characterized in that The step of redrawing the denoising feature set of the current round based on the feature similarity between each denoising feature of the current round and the reference image feature in the denoising feature set of the current round, in combination with the configuration threshold of the current round obtained from the threshold mapping relationship, to obtain the intermediate feature set of the current round, includes: Obtaining the feature similarity between each denoising feature of this round and the reference image feature in the denoising feature set of this round; Redraw the denoising features of this round whose feature similarity is greater than the configuration threshold of this round to obtain the intermediate features of this round; The denoising features of this round whose feature similarity is not greater than the configuration threshold of this round are used as the intermediate features of this round, and the obtained intermediate features of this round are arranged to obtain the intermediate feature set of this round.

6. The method according to claim 1, wherein The configuration parameters include configuration similarity, which is used to limit the participation of the current denoising feature set by limiting the lower limit of the similarity between any feature element in the current denoising feature set generated during the denoising and redrawing process and any feature element in the reference image features; the mapping relationship includes: a similarity mapping relationship between the noise ratio and the configuration similarity; The step of performing denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the at least one current round configuration parameter to obtain the current round intermediate feature set includes: Based on the content semantic features, denoising the previous intermediate feature set to obtain a denoising feature set for this round; Obtaining the element similarity between each feature element in each current round of denoising features and each feature element in the reference image features; Based on the obtained similarities of each element and the current round configuration similarity obtained from the similarity mapping relationship, the current round denoising feature set is redrawn to obtain the current round intermediate feature set.

7. The method according to claim 6, characterized in that The process of redrawing the denoising feature set of this round based on the obtained similarity of each element and the configuration similarity of this round obtained from the similarity mapping relationship to obtain the intermediate feature set of this round includes: For each current-round denoising feature in the current-round denoising feature set, perform the following operations: Based on the obtained similarities of each element, determining the number of elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity; The denoising features of this round are redrawn based on the number of elements to obtain intermediate features of this round; and the intermediate features of this round obtained are arranged to obtain an intermediate feature set of this round.

8. The method according to claim 7, characterized in that The determining, based on the obtained similarities of each element, the number of elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity includes: For each feature element in the current round of denoising features, the following operations are performed: based on the element similarity between a feature element in the current round of denoising features and each feature element in the reference image features, a feature element in the reference image features whose element similarity satisfies a similarity condition is used as a target element corresponding to the feature element; Based on the element correlation between each feature element in the current round of denoising features and its corresponding target element, the number of feature elements in the current round of denoising features whose element similarity is greater than the current round of configuration similarity is determined.

9. The method according to claim 6, characterized in that The process of redrawing the denoising feature set of this round based on the obtained similarity of each element and the configuration similarity of this round obtained from the similarity mapping relationship to obtain the intermediate feature set of this round includes: For each current-round denoising feature in the current-round denoising feature set, performing the following steps: determining, based on the obtained element similarities, a feature element in the current-round denoising feature whose element similarity is greater than the current-round configuration similarity as a to-be-redrawn element; and determining, in the reference image feature, a feature element whose element similarity to the to-be-redrawn element is greater than the current-round configuration similarity as a reference element; Based on the obtained element differences between each to-be-redrawn element and its corresponding reference element, each to-be-redrawn element in the current round of denoising features is adjusted to obtain the current round of intermediate features.

10. The method according to claim 9, characterized in that The configuration parameters include a configuration difference; the configuration difference is used to limit the participation of the denoising features in this round by limiting the difference between the denoising features in this round before and after redrawing; the mapping relationship includes: a difference mapping relationship between the noise ratio and the configuration difference; The step of adjusting the elements to be redrawn in the current round denoising features based on the obtained element differences between the elements to be redrawn and their corresponding reference elements to obtain the current round intermediate features includes: Based on the configuration difference of this round obtained from the difference mapping relationship, each of the elements to be redrawn and its corresponding element difference are fused in the denoising features of this round to obtain the intermediate features of this round.

11. The method according to any one of claims 1 to 10, characterized in that The obtaining of the content semantic features of the prompt content includes: When the prompt content is in text form, performing text feature extraction on the prompt content, and using the obtained text semantic features as content semantic features; When the prompt content is in the form of an image, performing image feature extraction on the prompt content, and using the obtained image semantic features as content semantic features; When the prompt content is in audio form, audio features are extracted from the prompt content, and the obtained audio semantic features are used as content semantic features.

12. An image synthesis device, characterized in that: include: Acquisition module: used for acquiring content semantic features of prompt content and acquiring an original image feature set including original image features of at least one noise image; A processing module is configured to perform multiple rounds of denoising and redrawing on the original image feature set in an iterative manner based on the content semantic features to obtain a synthetic image set; wherein each round of denoising and redrawing includes: The processing module is specifically configured to: obtain, based on a mapping relationship between at least one preset configuration parameter and the noise ratio, at least one configuration parameter of the current round corresponding to the noise ratio of the previous intermediate feature set obtained in the previous round of denoising and redrawing; each configuration parameter is negatively correlated with the noise ratio and is used to limit the degree of participation of a feature in denoising and redrawing; The processing module is specifically configured to: perform denoising and redrawing on the previous intermediate feature set based on the content semantic features according to the at least one current round configuration parameter to obtain the current round intermediate feature set; The processing module is specifically configured to use the obtained intermediate feature set of this round as the previous intermediate feature set to enter the next round of denoising and redrawing, until the final round of denoising and redrawing, and generate the synthetic image set based on the obtained intermediate feature set of this round.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer device, characterized in that: include: memory for storing computer programs; The processor is configured to call the computer program stored in the memory, and execute the method according to any one of claims 1 to 11 according to the obtained computer program.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1 to 11.