Image synthesis model training method and image synthesis method
By extracting and combining specific components in the pre-trained literary and genomic graph model, the image synthesis model is solved, and the problem of intricate handling of two-person photo details in the prior art is achieved, and high-quality two-person photo generation is achieved.
Patent Information
- Application Number
- CN202510161823.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, when generating a pair of photos, it is difficult to fine-tune the details such as glasses and hairstyles, resulting in low realism of details and affecting the group photo effect.
By extracting the text encoder, the first fusion component and the image generator in the pre-trained literary graph model, combining the feature extractor and the second fusion component, the image synthesis model is trained to generate a high-quality two-person photo.
It realizes the generation of high-quality photos of two faces in the scene described in the text, improving the authenticity of detailed features and group photo effects, and meeting the needs of users to take photos with characters in any scene.
Smart Images

Figure CN120182433A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to an image synthesis model training method and an image synthesis method. Background Art
[0002] With the development of artificial intelligence and computer vision technologies, related technologies have been able to generate group photos of two people. However, although related technologies have achieved the effect of group photos of two people to a certain extent, they still rely on a large amount of manual operations by users, and the usage scenarios are limited, and the requirements for lighting conditions are high.
[0003] Moreover, when related technologies are used to take group photos of two people, it is difficult to perform refined processing on detailed features such as glasses and hairstyles, resulting in low authenticity of details such as glasses and hairstyles, thus affecting the photo effect. Therefore, related technologies are difficult to generate high-quality group photos of two people. Summary of the Invention
[0004] The present disclosure provides an image synthesis model training method and an image synthesis method to solve at least one problem in related technologies. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, an image synthesis model training method is provided, including:
[0006] Extracting a text encoder, a first fusion component, and an image generator from a pre-trained text-to-image model;
[0007] Obtaining a sample image and a sample text, where the sample image is a group photo of a first sample face and a second sample face in a corresponding sample scenario, and the sample text is a text description of the sample scenario;
[0008] Inputting the sample text into the text encoder for text encoding to obtain sample text features;
[0009] Inputting the first sample face and the second sample face into a feature extractor for feature extraction to obtain a first sample face feature and a second sample face feature;
[0010] By respectively inputting the first sample face feature and the second sample face feature into a second fusion component for feature fusion, and inputting the sample text features into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text features, the first sample face feature, and the second sample face feature;
[0011] Based on the sample image and the predicted image, train the second fusion component and the feature extractor to obtain the trained second fusion component and the trained feature extractor;
[0012] Combine the trained second fusion component, the trained feature extractor, the text encoder, the first fusion component, and the image generator to obtain an image synthesis model.
[0013] In an exemplary embodiment, the feature extractor includes a pre-trained face encoder and a mapping component, and the mapping component is used to map the information in the image feature space to the text feature space; the step of inputting the first sample face and the second sample face into the feature extractor respectively for feature extraction to obtain the first sample face feature and the second sample face feature includes:
[0014] Input the first sample face and the second sample face into the face encoder respectively for face encoding to obtain the first sample face encoding and the second sample face encoding;
[0015] Input the first sample face encoding and the second sample face encoding into the mapping component respectively for information mapping to obtain the first sample face feature and the second sample face feature;
[0016] The step of training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes:
[0017] With the parameters of the face encoder frozen, train the mapping component based on the sample image and the predicted image.
[0018] In an exemplary embodiment, the step of triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature, and the second sample face feature by inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion includes:
[0019] Input the first sample face feature and the second sample face feature into the second fusion component respectively for feature fusion to obtain the first fused face feature and the second fused face feature;
[0020] The second fusion component inputs the first fused face feature and the second fused face feature into at least one network layer of the image generator, and the first fusion component inputs the fusion result of the sample text feature into at least one network layer of the image generator, so that the image generator fuses the fusion result of the sample text feature, the first fused face feature and the second fused face feature and generates a predicted image.
[0021] In an exemplary embodiment, the image generator is a diffusion model, and the predicted image generated by the image generator includes image prediction results corresponding to each diffusion step in at least one diffusion step; training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes:
[0022] Based on the sample image, determine the reference image corresponding to each diffusion step in the at least one diffusion step;
[0023] For any diffusion step in the at least one diffusion step, train the second fusion component and the feature extractor based on the difference between the reference image corresponding to the diffusion step and the corresponding image prediction result to obtain the trained second fusion component and the trained feature extractor.
[0024] In an exemplary embodiment, by respectively inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature and the second sample face feature, includes:
[0025] Initialize the current diffusion step;
[0026] Based on the preset noise and the current diffusion step, perform image noise prediction by fusing the first sample face feature, the second sample face feature and the sample text feature to obtain the image prediction result corresponding to the current diffusion step;
[0027] Update the preset noise based on the difference between the preset noise and the image prediction result corresponding to the current diffusion step;
[0028] Update the current diffusion step number. If the updated current diffusion step number is less than the diffusion step number threshold, repeat the step of predicting image noise by fusing the first sample face feature, the second sample face feature, and the sample text feature based on the preset noise and the current diffusion step number to obtain the image prediction result corresponding to the current diffusion step number.
[0029] In an exemplary embodiment, the second fusion component has the same structure as the first fusion component, and the connection manner between the second fusion component and the image generator is the same as the connection manner between the first fusion component and the image generator.
[0030] In an exemplary embodiment, both the first fusion component and the second fusion component are cross-attention components, which are used to guide the image generator to perform feature fusion based on cross-attention.
[0031] According to a second aspect of the embodiments of the present disclosure, there is provided an image synthesis method, and the method includes:
[0032] Obtain a first target face, a second target face, and a target text, where the target text is a text description of a target scene;
[0033] Input the first target face, the second target face, and the target text into an image synthesis model to obtain a target image, where the target image is a group photo of the first target face and the second target face in the target scene;
[0034] Wherein, the image synthesis model is trained by the image synthesis model training method according to any one of the first aspects.
[0035] In an exemplary embodiment, the image synthesis model includes a face encoder, a mapping component, a text encoder, a first fusion component, a second fusion component, and an image generator; the inputting the first target face, the second target face, and the target text into the image synthesis model to obtain a target image includes:
[0036] Input the first target face and the second target face into the face encoder for face encoding to obtain a first target face encoding and a second target face encoding;
[0037] Map the first target face encoding to a first feature map; fuse the first feature map with a first mask to obtain the first target face feature; map the second target face encoding to a second feature map; fuse the second feature map with a second mask to obtain the second target face feature; wherein, the first mask is used to guide the image generator to synthesize the corresponding first target face in the left region of the target image, and the second mask is used to guide the image generator to synthesize the corresponding second target face in the right region of the target image;
[0038] Input the target text into the text encoder for text encoding to obtain a target text feature;
[0039] By inputting the first target face feature and the second target face feature into the second fusion component, and inputting the target text feature into the first fusion component, trigger the first fusion component and the second fusion component to guide the image generator to generate the target image.
[0040] According to the third aspect of the embodiments of the present disclosure, there is provided an image synthesis model training device, including:
[0041] A component extraction module, configured to extract a text encoder, a first fusion component, and an image generator from a pre-trained text-to-image model;
[0042] A sample acquisition module, configured to acquire a sample image and a sample text, where the sample image is a group photo of a first sample face and a second sample face in a corresponding sample scene, and the sample text is a text description of the sample scene;
[0043] A training module, configured to perform:
[0044] Input the sample text into the text encoder for text encoding to obtain a sample text feature;
[0045] Input the first sample face and the second sample face into a feature extractor for feature extraction respectively to obtain a first sample face feature and a second sample face feature;
[0046] By inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion, trigger the first fusion component and the second fusion component to guide the image generator to generate a prediction image based on the feature fusion result of the sample text feature, the first sample face feature, and the second sample face feature;
[0047] Based on the sample image and the predicted image, train the second fusion component and the feature extractor to obtain the trained second fusion component and the trained feature extractor;
[0048] Combine the trained second fusion component, the trained feature extractor, the text encoder, the first fusion component, and the image generator to obtain an image synthesis model.
[0049] In an exemplary embodiment, the feature extractor includes a pre-trained face encoder and a mapping component, and the mapping component is used to map the information in the image feature space to the text feature space; the training module is configured to execute:
[0050] Input the first sample face and the second sample face into the face encoder for face encoding respectively to obtain the first sample face encoding and the second sample face encoding;
[0051] Input the first sample face encoding and the second sample face encoding into the mapping component for information mapping respectively to obtain the first sample face feature and the second sample face feature;
[0052] The training of the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes:
[0053] While freezing the parameters of the face encoder, train the mapping component based on the sample image and the predicted image.
[0054] In an exemplary embodiment, the training module is configured to execute:
[0055] Input the first sample face feature and the second sample face feature into the second fusion component for feature fusion respectively to obtain the first fused face feature and the second fused face feature;
[0056] The second fusion component inputs the first fused face feature and the second fused face feature into at least one network layer of the image generator, and the first fusion component inputs the fusion result of the sample text feature into at least one network layer of the image generator, so that the image generator fuses the fusion result of the sample text feature, the first fused face feature, and the second fused face feature and generates a predicted image.
[0057] In an exemplary embodiment, the image generator is a diffusion model, and the predicted image generated by the image generator includes the image prediction results corresponding to each diffusion step in at least one diffusion step; the training module is configured to execute:
[0058] Based on the sample image, determine a reference image corresponding to each diffusion step number among the at least one diffusion step number;
[0059] For any one of the at least one diffusion step number, based on the difference between the reference image corresponding to the diffusion step number and the corresponding image prediction result, train the second fusion component and the feature extractor to obtain a trained second fusion component and a trained feature extractor.
[0060] In an exemplary embodiment, the training module is configured to execute:
[0061] Initialize the current diffusion step number;
[0062] Based on a preset noise and the current diffusion step number, perform image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature to obtain an image prediction result corresponding to the current diffusion step number;
[0063] Update the preset noise based on the difference between the preset noise and the image prediction result corresponding to the current diffusion step number;
[0064] Update the current diffusion step number, and in the case where the updated current diffusion step number is less than a diffusion step number threshold, repeatedly execute the step of performing image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature based on the preset noise and the current diffusion step number to obtain an image prediction result corresponding to the current diffusion step number.
[0065] In an exemplary embodiment, the second fusion component has the same structure as the first fusion component, and the connection manner between the second fusion component and the image generator is the same as the connection manner between the first fusion component and the image generator.
[0066] In an exemplary embodiment, the training module is configured to execute:
[0067] Both the first fusion component and the second fusion component are cross-attention components for guiding the image generator to perform cross-attention based feature fusion.
[0068] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image synthesis device, including:
[0069] A target data acquisition module configured to execute acquiring a first target face, a second target face, and a target text, where the target text is a text description of a target scene;
[0070] A synthesis module, configured to execute inputting the first target face, the second target face, and the target text into an image synthesis model to obtain a target image, where the target image is a group photo of the first target face and the second target face in the target scene;
[0071] Wherein, the image synthesis model is trained by the image synthesis model training method according to any one of the first aspects.
[0072] In an exemplary embodiment, the image synthesis model includes a face encoder, a mapping component, a text encoder, a first fusion component, a second fusion component, and an image generator; the synthesis module is configured to execute:
[0073] Input the first target face and the second target face into the face encoder for face encoding to obtain a first target face encoding and a second target face encoding;
[0074] Map the first target face encoding to a first feature map; fuse the first feature map with a first mask to obtain the first target face feature; map the second target face encoding to a second feature map; fuse the second feature map with a second mask to obtain the second target face feature; wherein, the first mask is used to guide the image generator to synthesize the corresponding first target face in the left region of the target image, and the second mask is used to guide the image generator to synthesize the corresponding second target face in the right region of the target image;
[0075] Input the target text into the text encoder for text encoding to obtain a target text feature;
[0076] By inputting the first target face feature and the second target face feature into the second fusion component, and inputting the target text feature into the first fusion component, trigger the first fusion component and the second fusion component to guide the image generator to generate the target image.
[0077] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0078] A processor;
[0079] A memory for storing instructions executable by the processor;
[0080] Wherein, the processor is configured to execute the instructions to implement the image synthesis model training method or the image synthesis method according to any of the above embodiments.
[0081] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer storage medium, when instructions in the computer storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the image synthesis model training method or the image synthesis method described in any of the above embodiments.
[0082] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the image synthesis model training method or the image synthesis method described in any of the above embodiments.
[0083] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0084] In the embodiments of the present disclosure, a text branch, namely a text encoder, a first fusion component, and an image generator, can be extracted from a pre-trained text-to-image model, i.e., a text-to-image architecture. The text branch provides the ability to generate images based on text information, and an additional image branch, namely a feature extractor and a second fusion component, is added. The first sample face and the second sample face in the sample image are both input into the image branch, and the sample text corresponding to the sample image is input into the text branch, and a predicted image can be generated. Based on the predicted image and the sample image, the parameters of the image branch can be adjusted while freezing the parameters of the text branch, and then an image synthesis model can be obtained. The image synthesis model has the ability to generate a group photo of two faces in the scene described by the text.
[0085] Specifically, the text encoder of the image synthesis model can encode the text description of any scene, so that a group photo can be generated for two faces in any scene, thus realizing the effect of a group photo of two faces, meeting the user's need for a group photo with people in any scene, and the group photo effect is vivid and realistic.
[0086] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] The accompanying drawings herein are incorporated into the specification and constitute a part of the present disclosure, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0088] Figure 1 is a flowchart of an image synthesis model training method shown according to an exemplary embodiment.
[0089] Figure 2 is a flowchart of a feature extraction method shown according to an exemplary embodiment.
[0090] Figure 3It is a schematic structural diagram of a method for training an image synthesis model shown according to an exemplary embodiment.
[0091] Figure 4 It is a schematic flowchart of a method for generating a predicted image shown according to an exemplary embodiment.
[0092] Figure 5 It is a schematic flowchart of an image synthesis method shown according to an exemplary embodiment Figure 1 。
[0093] Figure 6 It is a schematic flowchart of an image synthesis method shown according to an exemplary embodiment Figure 2 。
[0094] Figure 7 It is a block diagram of an apparatus for training an image synthesis model shown according to an exemplary embodiment.
[0095] Figure 8 It is a block diagram of an image synthesis apparatus shown according to an exemplary embodiment.
[0096] Figure 9 It is a structural block of a computer device shown according to an exemplary embodiment Figure 1 ;
[0097] Figure 10 It is a structural block of a computer device shown according to an exemplary embodiment Figure 2 。 Detailed implementation manners
[0098] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0099] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0101] Figure 1 It is a flowchart of an image synthesis model training method shown according to an exemplary embodiment. The image synthesis model training method can be applied to an electronic device, which can be implemented by a server or a terminal alone, or can be implemented by the cooperation of a terminal and a server. Among them, the terminal can be, but is not limited to, physical devices such as smart phones, tablets, laptops, desktop computers, smart speakers, smart wearable devices, digital assistants, augmented reality devices, virtual reality devices, etc., and can also include software such as application programs running in physical devices. The server can be, but is not limited to, an independent server, or can be a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud, cloud storage, network services, cloud communication, intermediate services, domain name services, security services, and big data and artificial intelligence platforms, etc. Refer to Figure 1 As shown, the method includes the following steps.
[0102] In step S101, a text encoder, a first fusion component, and an image generator are extracted from a pre-trained text-to-image model.
[0103] The present disclosure does not limit the pre-trained text-to-image model, which can be any artificial intelligence model with text-to-image capabilities. The artificial intelligence model at least includes a text encoder, a first fusion component, and an image generator, where the text encoder is used to encode text information to obtain corresponding text features. The text features are input into the first fusion component so that the first fusion component can guide the image generator to perform feature fusion based on the text features, and then generate an image, thereby realizing text-to-image.
[0104] In one embodiment, the pre-trained text-to-image model can be a multi-functional image generation model. It can generate images through text, or modify images according to text, providing a text-to-image architecture that includes several components. The text encoder, the first fusion component, and the image generator extracted in step S101 are the text branches of the architecture, which can realize the related functions of text-to-image, that is, be able to generate pictures according to the specified text.
[0105] In step S102, a sample image and a sample text are obtained. The sample image is a group photo of a first sample face and a second sample face in a corresponding sample scene, and the sample text is a text description of the sample scene.
[0106] A group photo of two people can be used as a sample image, and the background, scene, theme, etc. of the group photo can be described in text form to form a text description corresponding to the sample image, that is, the sample text.
[0107] In step S103, the sample text is input into the text encoder for text encoding to obtain sample text features.
[0108] The text encoder is a component already existing in the text-to-image model and can be directly extracted for use. Inputting the sample text into the text encoder can trigger the text encoder to encode the sample text, obtaining sample text features in the text feature space.
[0109] In step S104, the first sample face and the second sample face are respectively input into the feature extractor for feature extraction to obtain the first sample face features and the second sample face features.
[0110] The present disclosure does not limit the feature extractor, as long as it can achieve the technical purpose of extracting features from the first sample face and the second sample face. For example, convolutional neural networks and various image encoders in related technologies can be used as the feature extractor.
[0111] In step S105, by respectively inputting the first sample face features and the second sample face features into the second fusion component for feature fusion, and inputting the sample text features into the first fusion component for feature fusion, the first fusion component and the second fusion component are triggered to guide the image generator to generate a predicted image based on the feature fusion result of the sample text features, the first sample face features, and the second sample face features.
[0112] The first fusion component is an inherent component of the text branch of the text-to-image model and can guide the image generator to generate an image based on the sample text features. The second fusion component and the feature extractor are the image branches set by the present disclosure for image synthesis. The second fusion component can guide the image generator to generate an image based on the first sample face features and the second sample face features. Under the combined action of the first fusion component, the second fusion component, and the image generator, the sample text features, the first sample face features, and the second sample face features are fully fused, and then the predicted image is generated.
[0113] In one embodiment, both the first fusion component and the second fusion component are cross-attention components, which are used to guide the image generator to perform cross-attention based feature fusion. The first fusion component can guide the image generator to perform cross-attention fusion on the sample text features, and the first fusion component itself can also perform cross-attention fusion on the sample text features. The second fusion component can guide the image generator to perform cross-attention fusion on the first sample face features and the second sample face features, and the second fusion component itself can also perform cross-attention fusion on the first sample face features and the second sample face features. By making full use of cross-attention based feature fusion, the feature fusion effect is enhanced, and the quality of the synthesized image is improved.
[0114] The present disclosure does not limit the structure of the second fusion component itself and its connection manner with the image generator. In an exemplary embodiment, the second fusion component has the same structure as the first fusion component, and the connection manner between the second fusion component and the image generator is the same as the connection manner between the first fusion component and the image generator. The architecture of the pre-trained text-to-image model is a mature architecture, which has been proven to have excellent performance through a large number of experiments. The first fusion component corresponds to the text branch of the text-to-image model, and the text branch can make good use of the text features output by the text encoder to generate images. The structure of the second fusion component and its connection manner with the image generator both belong to the additionally added image branch. The structure and connection relationship between the image branch and the text branch are based on the same inventive concept, which can enable the image branch where the second fusion component is located to also have excellent performance, so that the image features output by it can also be well utilized to generate images.
[0115] In step S106, based on the sample image and the predicted image, the second fusion component and the feature extractor are trained to obtain the trained second fusion component and the trained feature extractor; the trained second fusion component, the trained feature extractor, the text encoder, the first fusion component, and the image generator are combined to obtain an image synthesis model.
[0116] During the process of training the second fusion component and the feature extractor, the parameters of the second fusion component and the feature extractor are adjusted. The present disclosure does not limit the parameter adjustment method. For example, the gradient descent method can be used. The parameter adjustment can be stopped when the number of parameter adjustment times reaches a preset number threshold, or when the difference between the sample image and the predicted image is less than a preset difference, to obtain an image synthesis model. The present disclosure does not set the preset number threshold and the preset difference, which can be freely selected according to the actual situation and do not constitute an implementation obstacle.
[0117] Embodiments of the present disclosure can extract a text branch, namely a text encoder, a first fusion component, and an image generator, from a pre-trained text-to-image model, i.e., a text-to-image architecture. This text branch provides the ability to generate images based on text information, and an additional image branch, namely a feature extractor and a second fusion component, is added. The first sample face and the second sample face in the sample image are both input into the image branch, and the sample text corresponding to the sample image is input into the text branch, so that a predicted image can be generated. Based on the predicted image and the sample image, the parameters of the image branch can be adjusted while freezing the parameters of the text branch, thereby obtaining an image synthesis model. This image synthesis model has the ability to generate a group photo of two faces in the scene described by the text.
[0118] Specifically, the text encoder of this image synthesis model can encode the text description of any scene, so that a group photo can be generated for two faces in any scene, thereby achieving the effect of a group photo of two faces, meeting the user's need for a group photo with people in any scene, and the group photo effect is vivid and lifelike. Even details such as glasses and hairstyles can be vividly generated.
[0119] In an exemplary embodiment, the feature extractor includes a pre-trained face encoder and a mapping component, and the mapping component is used to map the information in the image feature space to the text feature space. Please refer to Figure 2 , which is a flowchart of a feature extraction method shown according to an exemplary embodiment. The step of inputting the first sample face and the second sample face into the feature extractor respectively for feature extraction to obtain the first sample face feature and the second sample face feature includes:
[0120] In step S201, the first sample face and the second sample face are respectively input into the face encoder for face encoding to obtain the first sample face encoding and the second sample face encoding.
[0121] This face encoder can be a pre-trained encoder, and the first sample face encoding and the second sample face encoding can be directly obtained by using it. The present disclosure does not limit the face encoder.
[0122] In step S202, the first sample face encoding and the second sample face encoding are respectively input into the mapping component for information mapping to obtain the first sample face feature and the second sample face feature.
[0123] Since the image generator used in the present disclosure itself belongs to the text branch of the text-to-image model, during the pre-training stage of the text-to-image model, the image generator is trained to have the ability to generate images based on the features in the text feature space, that is, the image generator is good at processing the features in the text feature space. However, the first sample face encoding and the second sample face encoding are located in the image feature space. In order to enable the image generator to play a role, the present disclosure uses a mapping component to map the first sample face encoding and the second sample face encoding located in the image feature space to the text feature space, obtaining the first sample face feature and the second sample face feature. Of course, the present disclosure does not limit the structure of the mapping component. For example, it can be formed by a convolutional network.
[0124] The aforementioned image generator is used in the text-to-image model. The aforementioned image generator can generate images based on the text features output by the text encoder under the guidance of the first fusion component, and the reliability and accuracy of the generated results are very high. The present disclosure uses a high-performance face encoder and a mapping component to map the face information to the text feature space as well, so that when the image generator generates images, it not only considers the text features but also the mapping results of the face information in the text feature space. This makes the generated images not only have the scenes corresponding to the text features but also have the characters related to the face information, integrating multi-dimensional information such as text and face to generate a group photo with high reliability and high accuracy, thus fully meeting the user's expectations.
[0125] In an exemplary embodiment, training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes: training the mapping component based on the sample image and the predicted image with the parameters of the face encoder frozen. In this embodiment, without adjusting the parameters of the face encoder, the effect of tuning the parameters of the second fusion component is applied to the mapping component, thereby quickly training the mapping component and improving the training speed.
[0126] In an exemplary embodiment, by respectively inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature, and the second sample face feature, includes:
[0127] Input the first sample face feature and the second sample face feature into the second fusion component respectively for feature fusion to obtain a first fused face feature and a second fused face feature; the second fusion component inputs the first fused face feature and the second fused face feature into at least one network layer of the image generator, and the first fusion component inputs the fusion result of the sample text feature into at least one network layer of the image generator, so that the image generator fuses the fusion result of the sample text feature, the first fused face feature and the second fused face feature and generates a predicted image.
[0128] By inputting the first sample face feature, the second sample face feature, and the sample text feature into at least one network layer of the image generator, the image generator is guided to gradually fuse these features during the image generation process. Through multi-level fusion, the fusion effect of multiple types of information is improved, and the quality of the generated image is enhanced.
[0129] Figure 3 It is a schematic diagram of the architecture of a method for training an image synthesis model shown according to an exemplary embodiment. The image synthesis model includes a text encoder, a first fusion component, and an image generator from a text-to-image model, and also includes a face encoder, a mapping component, and a second fusion component provided in the present disclosure. Both the first mapping component and the second mapping component are cross-attention components.
[0130] The text encoder can encode the sample text to obtain the sample text feature in the text feature space.
[0131] The face encoder can encode the input first sample face and second sample face to obtain a first sample face encoding and a second sample face encoding, thereby realizing the retention of face features.
[0132] Both the first sample face encoding and the second sample face encoding are mapped to the text feature space by the mapping component to obtain a first face feature and a second face feature. The first face feature and the second face feature are in the same space as the sample text feature, so they can be fused with each other.
[0133] The first fusion component and the second fusion component gradually fuse the output features with the image generator respectively, so that the image generator gradually fuses the sample text feature, the first face feature and the second face feature, and gradually generates a predicted image. Based on the predicted image and the sample image, the mapping component and the second fusion component are adjusted, so that the image synthesis model gradually has the ability to generate a group photo desired by the user.
[0134] The image generator receives three inputs: sample text features, first face features, and second face features, so as to better generate a high-fidelity and high-resolution group photo of two people. The first fusion component and the second fusion component independently guide the features they output to interact with the fusion in each network layer of the image generator, so that different faces are independent in the finally generated group photo, and the finally generated group photo has the background, scene, or theme described by the sample text features. Figure 3 Two faces and the text "**** theme" are shown. The image synthesis model finally outputs the group photo result of these two faces under the "**** theme", and this group photo result is the group photo expected by the user.
[0135] In an exemplary embodiment, the image generator is a diffusion model, and the predicted image generated by the image generator includes image prediction results corresponding to each diffusion step in at least one diffusion step; training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes: determining a reference image corresponding to each diffusion step in the at least one diffusion step based on the sample image; for any diffusion step in the at least one diffusion step, training the second fusion component and the feature extractor based on the difference between the reference image corresponding to the diffusion step and the corresponding image prediction result to obtain the trained second fusion component and the trained feature extractor.
[0136] The present disclosure does not limit the structure of the diffusion model, which can refer to related technologies and does not constitute an implementation obstacle. For any diffusion step, the reference image corresponding to this diffusion step can be calculated based on the sample image, and this reference image can be understood as the "ground truth" feature used when generating the image prediction result at this diffusion step. The present disclosure does not limit and explain the specific algorithm for determining the reference image based on the diffusion step, because this specific algorithm can refer to related technologies and does not constitute an implementation obstacle to the present disclosure. The present disclosure can adjust parameters based on any diffusion step to improve the training speed. Of course, the present disclosure does not limit the quantization method of the difference between the reference image corresponding to the diffusion step and the corresponding image prediction result, which does not constitute an implementation obstacle. The difference between the reference image corresponding to the diffusion step and the corresponding image prediction result can be understood as the loss generated when generating the image prediction result at the diffusion step.
[0137] Figure 4It is a schematic flowchart of a predicted image generation method shown according to an exemplary embodiment. By respectively inputting the first sample face feature and the second sample face feature into a second fusion component for feature fusion, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature, and the second sample face feature, including:
[0138] In step S401, initialize the current diffusion step.
[0139] Exemplarily, when step S401 is executed for the first time, the current diffusion step can be set to 0, that is, the current diffusion step step = 0.
[0140] In step S402, based on a preset noise and the current diffusion step, perform image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature, to obtain an image prediction result corresponding to the current diffusion step.
[0141] The present disclosure does not limit the preset noise obtained for the first time. Exemplarily, when the preset noise is obtained for the first time, the preset noise can be set to random white noise. When step S402 is executed for the first time, the random white noise, step = 0, and the first sample face feature, the second sample face feature, and the sample text feature can be input into the diffusion model together. Of course, the present disclosure does not limit the structure of the diffusion model.
[0142] In step S403, update the preset noise based on the difference between the preset noise and the image prediction result corresponding to the current diffusion step.
[0143] When step = 0, the image prediction result corresponding to step = 0 is predicted. Then, the difference between the preset noise input at step = 0 and the output image prediction result can be calculated, and this difference is used as the new preset noise to participate in the next step, that is, as the preset noise at step = 1.
[0144] In step S404, update the current diffusion step. In the case where the updated current diffusion step is less than the diffusion step threshold, repeat the step of performing image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature based on the preset noise and the current diffusion step to obtain an image prediction result corresponding to the current diffusion step.
[0145] Exemplarily, the current diffusion step can perform an increment operation by 1. If step = 0, then step can be updated to 1, thereby starting the image noise prediction for the next step until there are corresponding image prediction results for each diffusion step less than the diffusion step threshold. The present disclosure does not limit the diffusion step threshold. For example, it can be limited to 100.
[0146] The diffusion model is used to gradually generate image prediction results corresponding to each diffusion step, thereby gradually optimizing the quality of the image prediction results until the image prediction result generated in the last diffusion step is obtained. This image prediction result can be used as the synthesized image finally output by the image generator. This process is a process of gradually reducing noise, and the finally generated synthesized image has good image quality.
[0147] In one exemplary embodiment, an image synthesis method is disclosed. Figure 5 It is a flowchart of an image synthesis method shown according to an exemplary embodiment Figure 1 . The image synthesis method includes:
[0148] In step S501, a first target face, a second target face, and a target text are obtained, where the target text is a text description of a target scene.
[0149] The target text is used to indicate information such as the background, theme, and scene used in the group photo. The first target face and the second target face are role materials for the group photo.
[0150] In step S502, the first target face, the second target face, and the target text are input into an image synthesis model to obtain a target image, where the target image is a group photo of the first target face and the second target face in the target scene; wherein, the image synthesis model is trained by the aforementioned image synthesis model training method.
[0151] The image synthesis model trained by the present disclosure can generate a group photo of the first target face and the second target face, and the group photo uses the theme, background, or scene described in the target text. The group photo has a strong sense of reality, high resolution, and high quality. Moreover, it does not require professional photographic equipment, has no strict requirements on lighting conditions, and does not rely on manual processing by users, and can achieve automated image synthesis.
[0152] Figure 6 It is a flowchart of an image synthesis method shown according to an exemplary embodiment Figure 2. In an exemplary embodiment, the image synthesis model includes a face encoder, a mapping component, a text encoder, a first fusion component, a second fusion component, and an image generator; the step of inputting the first target face, the second target face, and the target text into the image synthesis model to obtain a target image includes:
[0153] In step S601, the first target face and the second target face are input into the face encoder for face encoding to obtain a first target face encoding and a second target face encoding.
[0154] In step S602, the first target face encoding is mapped to a first feature map; the first feature map is fused with a first mask to obtain the first target face feature; the second target face encoding is mapped to a second feature map; the second feature map is fused with a second mask to obtain the second target face feature; wherein, the first mask is used to guide the image generator to synthesize the corresponding first target face in the left region of the target image, and the second mask is used to guide the image generator to synthesize the corresponding second target face in the right region of the target image.
[0155] The present disclosure does not limit the first mask and the second mask, as long as the first mask can be used to guide the image generator to synthesize the corresponding first target face in the left region of the target image, and the second mask can be used to guide the image generator to synthesize the corresponding second target face in the right region of the target image. For example, the first mask can mask the features corresponding to the right region of the target image in the first feature map, and the second mask can mask the features corresponding to the left region of the target image in the first feature map.
[0156] The present disclosure proposes a left - right layout assumption for a two - person group photo, that is, a two - person group photo is usually in a left - right layout, that is, the two faces in the photo are presented left - right distributed. Therefore, the present disclosure adds indication information of the left - right distribution during model inference by using the first mask and the second mask, so as to guide the model to generate the first target face on the left and the second target face on the right. This can not only maintain the independence of the two faces in the group photo, but also make the faces in the group photo coordinated and improve the realism of the group photo.
[0157] In step S603, the target text is input into the text encoder for text encoding to obtain a target text feature; by inputting the first target face feature and the second target face feature into the second fusion component, and inputting the target text feature into the first fusion component, the first fusion component and the second fusion component are triggered to guide the image generator to generate the target image.
[0158] The present disclosure processes facial features and text features separately, calculates their cross-attention based on their respective corresponding fusion components, so as to maintain the independence of two faces in the target image. The first fusion component and the second fusion component achieve multi-level feature fusion. Specifically, a multi-level decoupled cross-attention mechanism can be implemented, so as to achieve the balance of the backgrounds corresponding to the two faces and the text in the target image, and can seamlessly fuse different individuals in the same photo while maintaining the quality of natural images.
[0159] In the image synthesis method of the present disclosure, by inputting two images, the faces in the two images are retained, and the scene is described according to the input text, so as to meet the requirement of taking a group photo of two faces in the corresponding scene. This method can be widely applied to various fields that require image synthesis, such as personal entertainment, commercial advertising, film production, educational tools, etc. For example, if a user wants to take a photo with a star but actually cannot take a photo with the star in person, the user can use this method at this time. Just input the photo of the user and the star, and the text corresponding to the scene of the group photo, such as "taking a photo on the beach", into the image synthesis model trained by the present disclosure, and the image synthesis model can output the group photo of the user and the star on the beach and the group photo has a high degree of authenticity. The present disclosure can meet the user's need to take a group photo with any person in any scene, and significantly improve the user experience.
[0160] Figure 7 It is a block diagram of an image synthesis model training device shown according to an exemplary embodiment. Referring to Figure 7 , the device includes:
[0161] The component extraction module 701 is configured to extract a text encoder, a first fusion component, and an image generator in a pre-trained text-to-image model;
[0162] The sample acquisition module 702 is configured to acquire a sample image and a sample text, where the sample image is a group photo of a first sample face and a second sample face in a corresponding sample scene, and the sample text is a text description of the sample scene;
[0163] The training module 703 is configured to perform:
[0164] Input the sample text into the text encoder for text encoding to obtain sample text features;
[0165] Input the first sample face and the second sample face into a feature extractor for feature extraction to obtain a first sample face feature and a second sample face feature;
[0166] By separately inputting the first sample face feature and the second sample face feature into a second fusion component for feature fusion, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature, and the second sample face feature;
[0167] Based on the sample image and the predicted image, training the second fusion component and the feature extractor to obtain a trained second fusion component and a trained feature extractor;
[0168] Combining the trained second fusion component, the trained feature extractor, the text encoder, the first fusion component, and the image generator to obtain an image synthesis model.
[0169] In an exemplary embodiment, the feature extractor includes a pre-trained face encoder and a mapping component, and the mapping component is used to map information in the image feature space to the text feature space; the training module 703 is configured to execute:
[0170] Separately inputting the first sample face and the second sample face into the face encoder for face encoding to obtain a first sample face encoding and a second sample face encoding;
[0171] Separately inputting the first sample face encoding and the second sample face encoding into the mapping component for information mapping to obtain the first sample face feature and the second sample face feature;
[0172] The training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain a trained second fusion component and a trained feature extractor includes:
[0173] While freezing the parameters of the face encoder, training the mapping component based on the sample image and the predicted image.
[0174] In an exemplary embodiment, the training module 703 is configured to execute:
[0175] Separately inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion to obtain a first fused face feature and a second fused face feature;
[0176] The second fusion component inputs the first fused face feature and the second fused face feature into at least one network layer of the image generator, and the first fusion component inputs the fusion result of the sample text feature into at least one network layer of the image generator, so that the image generator fuses the fusion result of the sample text feature, the first fused face feature, and the second fused face feature and generates a predicted image.
[0177] In an exemplary embodiment, the image generator is a diffusion model, and the predicted image generated by the image generator includes image prediction results corresponding to each diffusion step in at least one diffusion step; the training module 703 is configured to perform:
[0178] Based on the sample image, determine a reference image corresponding to each diffusion step in the at least one diffusion step;
[0179] For any diffusion step in the at least one diffusion step, based on the difference between the reference image corresponding to the diffusion step and the corresponding image prediction result, train the second fusion component and the feature extractor to obtain a trained second fusion component and a trained feature extractor.
[0180] In an exemplary embodiment, the training module 703 is configured to perform:
[0181] Initialize the current diffusion step;
[0182] Based on the preset noise and the current diffusion step, perform image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature to obtain an image prediction result corresponding to the current diffusion step;
[0183] Update the preset noise based on the difference between the preset noise and the image prediction result corresponding to the current diffusion step;
[0184] Update the current diffusion step, and if the updated current diffusion step is less than the diffusion step threshold, repeat the step of performing image noise prediction by fusing the first sample face feature, the second sample face feature, and the sample text feature based on the preset noise and the current diffusion step to obtain an image prediction result corresponding to the current diffusion step.
[0185] In an exemplary embodiment, the second fusion component has the same structure as the first fusion component, and the connection manner between the second fusion component and the image generator is the same as the connection manner between the first fusion component and the image generator.
[0186] In an exemplary embodiment, the training module 703 is configured to perform:
[0187] Both the first fusion component and the second fusion component are cross-attention components, which are used to guide the image generator to perform cross-attention based feature fusion.
[0188] Regarding the device in the above embodiments, the specific manners of the steps have been described in detail in the embodiments of the foregoing method, and will not be elaborated herein.
[0189] Figure 8 is a block diagram of an image synthesis device shown according to an exemplary embodiment. Referring to Figure 8 , the device includes:
[0190] A target data acquisition module 801, configured to acquire a first target face, a second target face, and a target text, where the target text is a text description of a target scene;
[0191] A synthesis module 802, configured to input the first target face, the second target face, and the target text into an image synthesis model to obtain a target image, where the target image is a group photo of the first target face and the second target face in the target scene;
[0192] Wherein, the image synthesis model is trained by the image synthesis model training method according to any item in the first aspect.
[0193] In an exemplary embodiment, the image synthesis model includes a face encoder, a mapping component, a text encoder, a first fusion component, a second fusion component, and an image generator; the synthesis module 802 is configured to perform:
[0194] Input the first target face and the second target face into the face encoder for face encoding to obtain a first target face encoding and a second target face encoding;
[0195] Map the first target face encoding to a first feature map; fuse the first feature map with a first mask to obtain the first target face feature; map the second target face encoding to a second feature map; fuse the second feature map with a second mask to obtain the second target face feature; wherein, the first mask is used to guide the image generator to synthesize the corresponding first target face in the left region of the target image, and the second mask is used to guide the image generator to synthesize the corresponding second target face in the right region of the target image;
[0196] Input the target text into the text encoder for text encoding to obtain a target text feature;
[0197] By inputting the first target face feature and the second target face feature into the second fusion component, and inputting the target text feature into the first fusion component, the first fusion component and the second fusion component are triggered to guide the image generator to generate the target image.
[0198] Regarding the device in the above embodiments, the specific manners of the steps have been described in detail in the embodiments of the foregoing method, and will not be elaborated herein.
[0199] Please refer to Figure 9 , which shows the structural block of a computer device provided by an exemplary embodiment of the present disclosure Figure 1 . The computer device may be a terminal. The computer device is used to implement the image synthesis model training method or the image synthesis method provided in the above embodiments. Specifically:
[0200] Generally, the computer device 900 includes: a processor 901 and a memory 902.
[0201] The processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In an exemplary embodiment, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In an exemplary embodiment, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process calculation operations related to machine learning.
[0202] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In an exemplary embodiment, the non-transitory computer-readable storage media in the memory 902 is used to store at least one instruction, at least one program, code set or instruction set, and the above at least one instruction, at least one program, code set or instruction set is configured to be executed by one or more processors to implement the above image synthesis model training method or image synthesis method.
[0203] In an exemplary embodiment, the computer device 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902 and the peripheral device interface 903 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 904, a touch display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908 and a power supply 909.
[0204] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the computer device 900, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0205] Please refer to Figure 10 which shows the structural block of a computer device provided by another exemplary embodiment of the present disclosure Figure 2 . The computer device may be a server for executing the above image synthesis model training method or image synthesis method. Specifically:
[0206] The computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including a random access memory (RAM) 1002 and a read-only memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 also includes a basic input / output system (I / O (Input / Output) system) 1006 for helping to transfer information between various devices in the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014 and other program modules 1011.
[0207] The basic input / output system 1006 includes a display 1008 for displaying information and input devices 1009 such as a mouse, keyboard, etc. for user input of information. Both the display 1008 and the input devices 1009 are connected to the central processing unit 1001 through an input / output controller 1100 connected to the system bus 1005. The basic input / output system 1006 may also include an input / output controller 1100 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1100 also provides outputs to a display screen, printer, or other types of output devices.
[0208] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is to say, the mass storage device 1007 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0209] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art know that computer storage media is not limited to the above several types. The above-mentioned system memory 1004 and the mass storage device 1007 can be collectively referred to as memory.
[0210] According to various embodiments of the present disclosure, the computer device 1000 may also be run by a remote computer on the network through a network such as the Internet. That is, the computer device 1000 may be connected to the network 1012 through the network interface unit 1011 connected to the system bus 1005, or in other words, the network interface unit 1011 may also be used to connect to other types of networks or remote computer systems (not shown).
[0211] The above-mentioned memory further includes a computer program, which is stored in the memory and configured to be executed by one or more processors to implement the above-mentioned image synthesis model training method or image synthesis method.
[0212] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the above-mentioned storage medium. When the at least one instruction, the at least one program, the code set, or the instruction set is executed by a processor, the above-mentioned image synthesis model training method or image synthesis method is implemented.
[0213] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0214] In an exemplary embodiment, a computer-readable storage medium including program code is further provided, such as a memory including program code. The above-mentioned program code can be executed by a processor to complete the above-mentioned image synthesis model training method or image synthesis method. Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact-disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0215] In an exemplary embodiment, a computer program product is further provided, including a computer program, which implements the above-mentioned image synthesis model training method or image synthesis method when executed by a processor.
[0216] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0217] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for training an image synthesis model, characterized in that: include: Extract the text encoder, the first fusion component, and the image generator from the pre-trained text-graph model; Obtaining a sample image and a sample text, wherein the sample image is a photo of a first sample face and a second sample face in a corresponding sample scene, and the sample text is a text description of the sample scene; Inputting the sample text into the text encoder for text encoding to obtain sample text features; Inputting the first sample face and the second sample face into a feature extractor for feature extraction respectively, to obtain first sample face features and second sample face features; By inputting the first sample facial feature and the second sample facial feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample facial feature and the second sample facial feature; Based on the sample image and the predicted image, training the second fusion component and the feature extractor to obtain a trained second fusion component and a trained feature extractor; The trained second fusion component, the trained feature extractor, the text encoder, the first fusion component and the image generator are combined to obtain an image synthesis model.
2. The method according to claim 1, characterized in that The feature extractor includes a pre-trained face encoder and a mapping component, wherein the mapping component is used to map information in an image feature space to a text feature space; the first sample face and the second sample face are respectively input into the feature extractor for feature extraction to obtain first sample face features and second sample face features, including: Inputting the first sample face and the second sample face into the face encoder respectively for face encoding to obtain a first sample face code and a second sample face code; Inputting the first sample face code and the second sample face code into the mapping component for information mapping respectively, to obtain the first sample face feature and the second sample face feature; The step of training the second fusion component and the feature extractor based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor includes: The mapping component is trained based on the sample image and the predicted image while freezing the parameters of the face encoder.
3. The method according to claim 1 or 2, characterized in that: The method comprises: inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature and the second sample face feature, including: Inputting the first sample facial feature and the second sample facial feature into the second fusion component for feature fusion respectively to obtain a first fused facial feature and a second fused facial feature; The second fusion component inputs the first fused facial features and the second fused facial features into at least one network layer of the image generator, and the first fusion component inputs the fusion result of the sample text features into at least one network layer of the image generator, so that the image generator fuses the fusion result of the sample text features, the first fused facial features and the second fused facial features to generate a predicted image.
4. The method according to claim 1, characterized in that The image generator is a diffusion model, and the predicted image generated by the image generator includes an image prediction result corresponding to each diffusion step in at least one diffusion step; the second fusion component and the feature extractor are trained based on the sample image and the predicted image to obtain the trained second fusion component and the trained feature extractor, including: Based on the sample image, determining a reference image corresponding to each diffusion step in the at least one diffusion step; For any diffusion step of the at least one diffusion step, based on the difference between the reference image corresponding to the diffusion step and the corresponding image prediction result, the second fusion component and the feature extractor are trained to obtain the trained second fusion component and the trained feature extractor.
5. The method according to claim 4, characterized in that The method comprises: inputting the first sample face feature and the second sample face feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample face feature and the second sample face feature, including: Initialize the current diffusion steps; Based on the preset noise and the current number of diffusion steps, image noise prediction is performed by fusing the first sample face feature, the second sample face feature and the sample text feature to obtain an image prediction result corresponding to the current number of diffusion steps; Update the preset noise based on the difference between the preset noise and the image prediction result corresponding to the current diffusion step number; Update the current diffusion step number, and when the updated current diffusion step number is less than the diffusion step number threshold, repeat the step of performing image noise prediction based on the preset noise and the current diffusion step number by fusing the first sample face feature, the second sample face feature and the sample text feature to obtain the image prediction result corresponding to the current diffusion step number.
6. The method according to claim 1, characterized in that The second fusion component has the same structure as the first fusion component, and the connection method between the second fusion component and the image generator is the same as the connection method between the first fusion component and the image generator.
7. The method according to claim 1, characterized in that The first fusion component and the second fusion component are both cross-attention components, which are used to guide the image generator to perform feature fusion based on cross-attention.
8. An image synthesis method, characterized in that: include: Acquire a first target face, a second target face, and a target text, wherein the target text is a text description of a target scene; Inputting the first target face, the second target face and the target text into an image synthesis model to obtain a target image, where the target image is a photo of the first target face and the second target face in the target scene; Wherein, the image synthesis model is trained by the image synthesis model training method described in any one of claims 1 to 7.
9. The method according to claim 8, characterized in that The image synthesis model includes a face encoder, a mapping component, a text encoder, a first fusion component, a second fusion component and an image generator; the first target face, the second target face and the target text are input into the image synthesis model to obtain a target image, including: Inputting the first target face and the second target face into the face encoder for face encoding to obtain a first target face encoding and a second target face encoding; Mapping the first target face encoding to a first feature map; fusing the first feature map with a first mask to obtain the first target face feature; mapping the second target face encoding to a second feature map; fusing the second feature map with a second mask to obtain the second target face feature; wherein the first mask is used to guide the image generator to synthesize the corresponding first target face in the left area of the target image, and the second mask is used to guide the image generator to synthesize the corresponding second target face in the right area of the target image; Inputting the target text into the text encoder for text encoding to obtain target text features; By inputting the first target facial feature and the second target facial feature into the second fusion component, and inputting the target text feature into the first fusion component, the first fusion component and the second fusion component are triggered to guide the image generator to generate the target image.
10. An image synthesis model training device, characterized in that: include: A component extraction module is configured to extract a text encoder, a first fusion component, and an image generator in a pre-trained text-graph model; A sample acquisition module is configured to acquire a sample image and a sample text, wherein the sample image is a photo of a first sample face and a second sample face in a corresponding sample scene, and the sample text is a text description of the sample scene; The training module is configured to execute: Inputting the sample text into the text encoder for text encoding to obtain sample text features; Inputting the first sample face and the second sample face into a feature extractor for feature extraction respectively, to obtain first sample face features and second sample face features; By inputting the first sample facial feature and the second sample facial feature into the second fusion component for feature fusion respectively, and inputting the sample text feature into the first fusion component for feature fusion, triggering the first fusion component and the second fusion component to guide the image generator to generate a predicted image based on the feature fusion result of the sample text feature, the first sample facial feature and the second sample facial feature; Based on the sample image and the predicted image, training the second fusion component and the feature extractor to obtain a trained second fusion component and a trained feature extractor; The trained second fusion component, the trained feature extractor, the text encoder, the first fusion component and the image generator are combined to obtain an image synthesis model.
11. An image synthesis device, characterized in that: include: A target data acquisition module is configured to acquire a first target face, a second target face, and a target text, wherein the target text is a text description of a target scene; a synthesis module, configured to input the first target face, the second target face and the target text into an image synthesis model to obtain a target image, wherein the target image is a photo of the first target face and the second target face in the target scene; Wherein, the image synthesis model is trained by the image synthesis model training method described in any one of claims 1 to 7.
12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image synthesis model training method as described in any one of claims 1-7, or the image synthesis method as described in claim 8 or 9.
13. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device executes the image synthesis model training method as described in any one of claims 1-7, or the image synthesis method as described in any one of claims 8 or 9.
14. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the image synthesis model training method as described in any one of claims 1 to 7, or the image synthesis method as described in claim 8 or 9.