Image synthesis model generation method, device, electronic device and medium

By training the target synthesizer and encoder, using a small number of sample images and features to generate an image synthesis model, the problem of high generation cost of image synthesis model is solved, and the cost reduction of the model and task decoupling and multiplexing are achieved.

CN116152131BActive Publication Date: 2025-09-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310102206.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-29
Publication Date
2025-09-02
Estimated Expiration
2043-01-29

AI Technical Summary

Technical Problem

In the prior art, the generation cost of image synthesis models is relatively high, mainly due to the difficulty and large number of paired data acquisition.

Method used

By training the target synthesizer and the encoder to be trained based on the first sample image, an image synthesis model matching the feature type is generated, and the encoder is trained using a small number of second sample images and input features to reduce the need for paired data.

Benefits of technology

This reduces the number of paired sample data during the image synthesis model generation process, reduces the cost, and realizes the reusability of the synthesizer in different synthesis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152131B_ABST
    Figure CN116152131B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, electronic device, and medium for generating an image synthesis model, and relates to the field of computer technology. The method trains a synthesizer to be trained based on a first sample image in response to a model generation instruction for a first synthesis task to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoded features of the image; a reference image corresponding to the first input feature is generated using the target synthesizer and an encoder to be trained; the encoder to be trained is trained based on the reference image and a second sample image to generate a first encoder matching the feature type of the first synthesis task; and the first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task. In this way, the amount of paired sample data required to generate the image synthesis model can be reduced to a certain extent, thereby lowering the cost of generating the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device, and medium for generating an image synthesis model. Background Art

[0002] With the development of network technology, image synthesis technology has been applied to more and more fields. For example, human body image synthesis has a wide range of applications in the field of virtual try-on. Through image synthesis, high-definition images that are difficult to distinguish between true and false can be synthesized.

[0003] In the existing technology, a large amount of paired data is often used as training samples to train the generative model as a whole to obtain an image synthesis model. However, it is difficult to obtain paired data. Since this method requires a large number of paired samples, the generation cost of the image synthesis model is high. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, and medium for generating an image synthesis model to at least solve the problem of how to reduce the generation cost of the image synthesis model. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a method for generating an image synthesis model is provided, comprising:

[0006] In response to a model generation instruction for a first synthesis task, training a synthesizer to be trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoding features of the image;

[0007] generating, by the target synthesizer and the encoder to be trained, a reference image corresponding to a first input feature, wherein the first input feature is extracted from a second sample image according to a feature type of the first synthesis task, and the number of the second sample images is less than the number of the first sample images;

[0008] Based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate encoding features of the image based on the first input features;

[0009] The first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

[0010] Optionally, generating a control image corresponding to the first input feature by using the target synthesizer and the encoder to be trained includes:

[0011] Inputting the first input feature into the encoder to be trained to obtain an encoding feature of the second sample image corresponding to the first input feature;

[0012] inputting the encoded features of the second sample image into the target synthesizer to generate a control image;

[0013] The training of the encoder to be trained based on the control image and the second sample image includes: adjusting the model parameters in the encoder to be trained based on the control image and the second sample image, and continuing training based on the adjusted encoder to be trained until a first preset stop condition is reached, and determining the encoder to be trained at the end of training as the first encoder.

[0014] Optionally, the training the synthesizer to be trained based on the first sample image to generate a target synthesizer includes:

[0015] Inputting the preset sample coding features into the synthesizer to be trained, performing image synthesis processing, and obtaining a first image;

[0016] Based on the first image and the first sample image, the model parameters in the synthesizer to be trained are adjusted, and training is continued based on the adjusted synthesizer to be trained until a second preset stop condition is reached, and the synthesizer to be trained at the end of training is determined as the target synthesizer.

[0017] Optionally, the method further includes:

[0018] In response to a model generation instruction for a second synthesis task, obtaining a third sample image and a second input feature; the second input feature is extracted from the third sample image according to a feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task;

[0019] Based on the target synthesizer, the third sample image, and the second input feature, the encoder to be trained is trained to generate a second encoder that matches the feature type of the second synthesis task; the second encoder is used to generate encoding features of the image based on the second input feature;

[0020] The second encoder is combined with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task.

[0021] Optionally, the method further includes:

[0022] Based on the feature type of the first synthesis task, selecting a feature extraction model corresponding to the feature type from preset feature extraction models as a target feature extraction model;

[0023] Features matching the feature type of the first synthesis task are extracted from the second sample image based on the target feature extraction model to obtain the first input features.

[0024] Optionally, the designated type of image includes a human body image; the feature type includes at least one of a two-dimensional key point of a human body, a human body segmentation map, and a three-dimensional model of a human body.

[0025] According to a second aspect of an embodiment of the present disclosure, there is provided an image synthesis method, comprising:

[0026] Based on a target synthesis task, obtaining a target input feature of the target synthesis task and a target feature type corresponding to the target input feature;

[0027] Based on the target feature type, selecting an image synthesis model whose encoder matches the target feature type from pre-trained image synthesis models as the target image synthesis model;

[0028] The target input feature is used as the input of the target image synthesis model to obtain a target image synthesized by the target image synthesis model based on the target input feature.

[0029] According to a third aspect of an embodiment of the present disclosure, there is provided an image synthesis model generation device, comprising:

[0030] a synthesizer training module configured to execute, in response to a model generation instruction for a first synthesis task, training a synthesizer to be trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoded features of the image;

[0031] a reference image generation module configured to generate, through the target synthesizer and the encoder to be trained, a reference image corresponding to a first input feature; the first input feature is extracted from a second sample image according to a feature type of the first synthesis task; and the number of the second sample images is less than the number of the first sample images;

[0032] a first encoder training module configured to train the encoder to be trained based on the control image and the second sample image to generate a first encoder matching the feature type of the first synthesis task; the first encoder is configured to generate encoding features of an image based on the first input features;

[0033] A first combination module is configured to combine the first encoder with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

[0034] Optionally, the control image generation module includes:

[0035] a feature input submodule, configured to input the first input feature into the encoder to be trained to obtain an encoding feature of the second sample image corresponding to the first input feature;

[0036] The control image generation module is specifically configured to input the encoding features of the second sample image into the target synthesizer to generate a control image;

[0037] The first encoder training module is specifically configured to adjust the model parameters in the encoder to be trained based on the control image and the second sample image, and continue training based on the adjusted encoder to be trained until a first preset stop condition is reached, and the encoder to be trained at the end of training is determined to be the first encoder.

[0038] Optionally, the synthesizer training module includes:

[0039] a sample feature input submodule configured to input a preset sample coding feature into the synthesizer to be trained, perform image synthesis processing, and obtain a first image;

[0040] The synthesizer parameter adjustment sub-submodule is configured to adjust the model parameters in the synthesizer to be trained based on the first image and the first sample image, and continue training based on the adjusted synthesizer to be trained until a second preset stop condition is reached, and determine the synthesizer to be trained at the end of training as the target synthesizer.

[0041] Optionally, the device further includes:

[0042] an acquisition module configured to execute, in response to a model generation instruction for a second synthesis task, acquire a third sample image and a second input feature; the second input feature is extracted from the third sample image according to a feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task;

[0043] a second encoder training module configured to train the encoder to be trained based on the target synthesizer, the third sample image, and the second input features to generate a second encoder matching the feature type of the second synthesis task; the second encoder is configured to generate encoding features of an image based on the second input features;

[0044] The second combination module is configured to combine the second encoder with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task.

[0045] Optionally, the device further includes:

[0046] A selection module is configured to execute, based on the feature type of the first synthesis task, selecting a feature extraction model corresponding to the feature type from preset feature extraction models as a target feature extraction model;

[0047] The feature extraction module is configured to extract features matching the feature type of the first synthesis task from the second sample image based on the target feature extraction model to obtain the first input feature.

[0048] Optionally, the designated type of image includes a human body image; the feature type includes at least one of a two-dimensional key point of a human body, a human body segmentation map, and a three-dimensional model of a human body.

[0049] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0050] processor;

[0051] a memory for storing instructions executable by the processor;

[0052] The processor is configured to execute the instructions to implement the method as described in any one of the first aspect or the second aspect.

[0053] According to a fifth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device executes a method as described in any one of the first aspect or the second aspect.

[0054] According to a sixth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes readable program instructions. When the readable program instructions are executed by a processor of an electronic device, the electronic device executes the method as described in any one of the first aspect or the second aspect.

[0055] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: In the embodiments of the present disclosure, by responding to the model generation instruction for the first synthesis task, the synthesizer to be trained is trained based on the first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoding features of the image; a control image corresponding to the first input feature is generated through the target synthesizer and the encoder to be trained; the first input feature is extracted from the second sample image according to the feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images; based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate the encoding features of the image based on the first input feature; the first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task. In this way, since the synthesizer is independent of the input type of the synthesis task, the synthesizer only needs to synthesize the image based on the encoding features of the image. Therefore, there is no need to use a large amount of paired data to train the synthesizer to be trained according to the synthesis task. It is only necessary to obtain a small amount of second sample images and first input features to train the encoder to be trained. To a certain extent, it can reduce the amount of paired sample data required in the process of generating the image synthesis model and reduce the generation cost of the model.

[0056] At the same time, by adaptively training the synthesizer and encoder separately, the synthesis process and the synthesis task are decoupled. When a new synthesis task is received, there is no need to retrain the synthesizer, only the encoder needs to be trained, thus achieving the reusability of the synthesizer for different synthesis tasks.

[0057] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0059] Figure 1 is a flowchart of a method for generating an image synthesis model according to an exemplary embodiment;

[0060] Figure 2 is a schematic diagram of a training process of an encoder to be trained according to an exemplary embodiment;

[0061] Figure 3is a schematic diagram of a training process of a synthesizer to be trained according to an exemplary embodiment;

[0062] Figure 4 is a flowchart of another method for generating an image synthesis model according to an exemplary embodiment;

[0063] Figure 5 is a schematic diagram showing an input type according to an exemplary embodiment;

[0064] Figure 6 is a schematic diagram showing image synthesis according to an exemplary embodiment;

[0065] Figure 7 is a flowchart of an image synthesis method according to an exemplary embodiment;

[0066] Figure 8 is a schematic diagram showing another image synthesis method according to an exemplary embodiment;

[0067] Figure 9 is a block diagram of a device for generating an image synthesis model according to an exemplary embodiment;

[0068] Figure 10 is a block diagram of an image synthesis device according to an exemplary embodiment;

[0069] Figure 11 is a block diagram of a device for generating an image synthesis model according to an exemplary embodiment;

[0070] Figure 12 It is a block diagram of another apparatus for generating an image synthesis model according to an exemplary embodiment. DETAILED DESCRIPTION

[0071] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0072] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0073] Figure 1 is a flow chart of a method for generating an image synthesis model according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps:

[0074] Step 101: In response to a model generation instruction for a first synthesis task, a synthesizer to be trained is trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on encoding features of the image.

[0075] Wherein, the above-mentioned first synthesis task is generated according to actual needs, which may include the input features of the synthesis task, and different input features may characterize the input feature type of the synthesis task and the type of image to be synthesized. Wherein, the above-mentioned first sample image refers to the sample image used to train the synthesizer to be trained, which may be a sample image set, which includes a large number of pre-acquired images. Wherein, the above-mentioned specified type refers to the type of image to be synthesized, for example, it may be a human body image, a human face image or an animal image, etc., which may be set according to actual needs. Accordingly, the above-mentioned first sample image is consistent with the specified type of image. For example, when it is necessary to synthesize a human body image, a large number of videos and images containing human body movements can be downloaded from the Internet in advance, and a large number of human body images can be obtained as the first sample image by video frame extraction.

[0076] The synthesizer to be trained may be pre-built or pre-selected. Specifically, a previously built synthesizer may be randomly selected as the synthesizer to be trained, or a synthesizer with poor synthesis effect may be selected as the synthesizer to be trained, etc. Specifically, the goal of training the synthesizer to be trained is to learn the distribution of a specified type of image so that it can synthesize a high-quality specified type of image based on the input vector information. The encoding features of the image may conform to specified parameters, and the specified parameters may include the distribution type and dimension. For example, it may be a vector feature with a dimension of 128 that conforms to a Gaussian distribution.

[0077] Specifically, the synthesizer to be trained can be constructed through an image synthesis network, for example, a generative adversarial network or a diffusion model, etc., and the embodiments of the present disclosure are not limited to this.

[0078] Step 102: Generate a control image corresponding to a first input feature through the target synthesizer and the encoder to be trained; the first input feature is extracted from a second sample image according to the feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images.

[0079] Step 103: Based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate encoding features of the image based on the first input features.

[0080] The second sample images are sample images used to train the encoder to be trained, and may also be a set of sample images, which may be randomly selected from the first sample images, thereby eliminating the need to obtain sample images again. It should be noted that since the encoder only needs to encode the input features to generate the corresponding image encoding features, the encoder has a simpler network structure and fewer parameters than a synthesizer that needs to synthesize images. Therefore, the number of the second sample images is often less than that of the first sample images.

[0081] Among them, the above-mentioned first synthesis task is generated according to the actual needs of the user, which may include input information, so the above-mentioned feature type refers to the type of input information. It should be noted that image synthesis refers to synthesizing an image that matches the input information based on the input information features, and different synthesis tasks created by different users according to actual needs, the input information features may also be of different types, for example, it can be two-dimensional information or three-dimensional information, and the two-dimensional information can further be in the form of key points or segmentation maps. Therefore, in the embodiment of the present disclosure, the encoder to be trained can be trained in a targeted manner according to the feature type in the synthesis task.

[0082] It is understandable that different feature information can be extracted from the same sample image using different extraction methods or models. In the disclosed embodiment, the first input feature can be extracted from the second sample image according to the feature type of the first synthesis task. Furthermore, when there are multiple second sample images, the first input feature and the second sample image are in a one-to-one correspondence, thereby forming multiple pairs of sample data.

[0083] The goal of the encoder is to output appropriate image coding features based on the input features, so that the target synthesizer can synthesize an image that matches the synthesis task based on the image coding features. Specifically, the training method for the encoder to be trained can be to use the obtained first input features as input, obtain the corresponding image coding features through the encoder to be trained, and then input the vector information into the target synthesizer, so that the target synthesizer generates a reference image. Based on the difference between the reference image and the corresponding second sample image, the parameters of the encoder to be trained are adjusted and training continues until the difference between the synthesized reference image and the second sample image meets the requirements, thereby obtaining the first encoder.

[0084] Step 104: Combine the first encoder with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

[0085] Among them, the above-mentioned combination operation refers to connecting the first encoder and the target synthesizer in series, that is, using the output of the first encoder as the input of the target synthesizer. It can be understood that through the above-mentioned steps 101 to 103, a target synthesizer that can synthesize a specified type of image according to the image coding features, and a first encoder that matches the feature type of the first synthesis task and can generate image coding features according to the first input features can be obtained. Therefore, in the embodiment of the present disclosure, by combining the above-mentioned first encoder with the target synthesizer, a first image synthesis model that can synthesize a specified type of image according to the first input features can be obtained. Specifically, the above-mentioned first image synthesis model can receive input features that meet the feature type of the first synthesis task, generate image coding features based on the input features, and further synthesize images based on the image coding features.

[0086] Furthermore, since the target synthesizer in the embodiment of the present disclosure is independent of the feature type of the synthesis task, the target synthesizer can be combined with different encoders to obtain image synthesis models for different feature types. Therefore, when a new synthesis task is received, there is no need to retrain the synthesizer, only the encoder needs to be trained.

[0087] In summary, the image synthesis model generation method provided by the embodiment of the present disclosure trains the synthesizer to be trained based on the first sample image in response to the model generation instruction for the first synthesis task to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoding features of the image; a control image corresponding to the first input feature is generated through the target synthesizer and the encoder to be trained; the first input feature is extracted from the second sample image according to the feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images; based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate the encoding features of the image based on the first input feature; the first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task. In this way, since the synthesizer is independent of the input type of the synthesis task, the synthesizer only needs to synthesize the image based on the encoding features of the image. Therefore, there is no need to use a large amount of paired data to train the synthesizer to be trained according to the synthesis task. It is only necessary to obtain a small amount of second sample images and first input features to train the encoder to be trained. To a certain extent, it can reduce the amount of paired sample data required in the process of generating the image synthesis model and reduce the generation cost of the model.

[0088] At the same time, by adaptively training the synthesizer and encoder separately, the synthesis process and the synthesis task are decoupled. When a new synthesis task is received, there is no need to retrain the synthesizer, only the encoder needs to be trained, thus achieving the reusability of the synthesizer for different synthesis tasks.

[0089] In an optional embodiment, the operation of generating a reference image corresponding to the first input feature by using the target synthesizer and the encoder to be trained may include the following steps:

[0090] Step 201: Input the first input feature into the encoder to be trained to obtain the encoding feature of the second sample image corresponding to the first input feature.

[0091] Step 202: Input the encoding features of the second sample image into the target synthesizer to generate a comparison image.

[0092] The operation of training the encoder to be trained based on the control image and the second sample image may specifically include the following steps in the embodiment of the present disclosure:

[0093] Step 203: Adjust the model parameters of the encoder to be trained based on the control image and the second sample image, and continue training based on the adjusted encoder to be trained until a first preset stop condition is reached, and determine the encoder to be trained at the end of training as the first encoder.

[0094] The goal of the encoder is to output corresponding image coding features based on the input features. These coding features enable the synthesizer to synthesize an image that matches the input features. Therefore, the image coding features output by the encoder often correspond to the input features. The above input features can be understood as control conditions, which include the conditions or properties that the final synthesized image must meet.

[0095] Specifically, the encoder to be trained can be an autoencoder (AE), which can compress input features into a latent space code (i.e., the image coding features described above). The latent space code can include attributes of the input features, and the dimensions and other information of the image coding features generated by AEs with different model parameters are also different. Therefore, the disclosed embodiments can dynamically adjust the model parameters of the encoder through training, so that it can generate image coding features that meet the requirements of the target synthesizer.

[0096] After the image coding features corresponding to the first input features are obtained by the encoder to be trained, they can be input into the target synthesizer obtained in step 101 to synthesize the corresponding image, that is, the reference image. Furthermore, based on the difference between the reference image and the second sample image, the model parameters in the encoder to be trained can be adjusted, and the above-mentioned training steps can be continued using the adjusted encoder to be trained until a first preset stopping condition is reached. When the first preset stopping condition is reached, the training is determined to be complete.

[0097] Among them, the above-mentioned first preset stop condition refers to the training stop condition of the encoder to be trained, which may be that the training time reaches the preset time, or the number of model parameter adjustments reaches the preset number of times, or the preset loss function value reaches the minimum, etc. The embodiment of the present disclosure does not limit this.

[0098] Specifically, taking the above-mentioned first preset stopping condition as the preset loss function value reaching the minimum as an example, the preset loss function can be the reconstruction loss L(.) between the reference image and the second sample image, for example, L1 or L2 loss, and then the loss function can be obtained in the following form:

[0099]

[0100] Wherein, the above G represents the target synthesizer, E represents the encoder to be trained, the above yi and Ii represent the input features and the second sample image input to the encoder to be trained respectively, wherein the above i is used to represent different sample images, and m represents the number of second sample images. Thus, the above E(yi) represents the coding features or latent space coding of the second sample image generated by the encoder to be trained, the above G(E(yi)) represents the reference image generated by the target synthesizer according to the input coding features, and L(G(E(yi)), Ii) represents the reconstruction loss between the reference image and the second sample image, which can represent the degree of difference between the reference image and the sample image. It can be understood that the smaller the difference between the reference image and the sample image, the better the image synthesis effect. Accordingly, the objective function in the training process of the encoder to be trained can be:

[0101]

[0102] The above-mentioned objective function can be the first preset stop condition set in the embodiment of the present disclosure. When the loss function value between the control image and the sample image reaches the minimum value, it indicates that the current encoder to be trained has met the training requirements. At this time, the training is completed, so that the current model parameters can be used as the model parameters of the final first encoder, and the current encoder to be trained is determined as the first encoder.

[0103] Figure 2 FIG. 1 is a schematic diagram of a training process of an encoder to be trained according to an exemplary embodiment. Figure 2 As shown in the figure, the input control condition in the figure refers to the first input feature mentioned above, that is, y in the above formulas (1) and (2) i , where the encoder refers to the encoder to be trained, the noise refers to the image coding features or latent space coding generated by the encoder according to the input control conditions, and the pre-trained generator refers to the target synthesizer. The pre-trained generator can generate a corresponding synthetic image through the input noise (coding features), so that the encoder to be trained can be trained through the reconstruction loss between the real image (second sample image) and the synthetic image.

[0104] In the disclosed embodiment, the encoding features of the second sample image corresponding to the first input features are obtained by inputting the first input features into the encoder to be trained; the encoding features of the second sample image are input into the target synthesizer to generate a reference image; the model parameters in the encoder to be trained are adjusted based on the reference image and the second sample image, and the training is continued based on the adjusted encoder to be trained until the first preset stop condition is reached, and the encoder to be trained at the end of the training is determined to be the first encoder. In this way, a reference image is generated for the encoding features output by the encoder to be trained by the trained target synthesizer, and the model parameters of the encoder to be trained can be adjusted by the reference image and the corresponding second sample image. The encoder to be trained can be trained from the perspective of image synthesis effect, and a first encoder that matches the feature type of the first synthesis task and meets the requirements of the image synthesis effect is obtained, thereby ensuring the image synthesis effect of the first image synthesis model obtained in the corresponding feature type.

[0105] In an optional embodiment, the operation of training the synthesizer to be trained based on the first sample image to generate the target synthesizer may include the following steps:

[0106] Step 301: Input the preset sample coding features into the synthesizer to be trained, perform image synthesis processing, and obtain a first image.

[0107] Among them, the above-mentioned sample coding features can be pre-selected. It should be noted that the purpose of training the synthesizer is to enable it to learn the distribution manifold of a specified type of image. Therefore, the above-mentioned sample vectors can be randomly sampled from noise that conforms to a known distribution (for example, a Gaussian distribution), so that a mapping between coding features that conform to a known distribution and real images can be established, so that image synthesis can be performed using different coding features. It can be seen that in the embodiment of the present disclosure, the training of the synthesizer to be trained does not require manual labeling of sample information, and only requires obtaining a large number of sample images and randomly sampled sample coding features for training, thereby achieving unsupervised training.

[0108] Step 302: Based on the first image and the first sample image, adjust the model parameters in the synthesizer to be trained, and continue training based on the adjusted synthesizer to be trained until a second preset stop condition is reached, and determine the synthesizer to be trained at the end of training as the target synthesizer.

[0109] Among them, the above-mentioned second preset stopping condition refers to the training stopping condition of the synthesizer to be trained, which may be that the training time reaches the preset time, or the number of model parameter adjustments reaches the preset number of times, or the preset loss function value reaches the threshold, etc. The embodiment of the present disclosure does not limit this. When the second preset stopping condition is reached, it is determined that the training of the synthesizer to be trained is completed.

[0110] Specifically, the synthesizer to be trained can be a generative model such as a generative adversarial network or a diffusion model. A generative adversarial network can synthesize images that are difficult to distinguish between real and fake through adversarial training, while a diffusion model defines a Markov chain that includes two steps: diffusion and inverse diffusion. The diffusion process gradually adds random noise to the data, while the inverse diffusion process recovers the data samples from the noise. The present disclosure embodiments can employ any generative model and are not limited thereto.

[0111] For example, taking the generative adversarial network as an example, Figure 3 FIG. 1 is a schematic diagram of a training process of a synthesizer to be trained according to an exemplary embodiment. Figure 3 As shown in Figure 1, a generative adversarial network (GAN) consists of a generator and a discriminator. Both learn through adversarial training. After training, the generator can synthesize images from noise. The input to the generator is noise z, a vector of a specified dimension, generally following a standard Gaussian distribution. The input to the discriminator is either a synthesized image or a real image, and the output is a true / false score, typically between 0 and 1. The objective function of a GAN can be expressed as follows:

[0112] min G max D E x~q(x) [logD(x)]+E z~p(z) [log(1-D(G(z)))]

[0113] Wherein, the above x represents the real image, that is, the first sample image, x~q(x) means that the real image obeys the q(x) distribution, q(x) represents the distribution obeyed by the real image, which is an unknown distribution. Correspondingly, z is the noise, that is, the above-mentioned preset sample encoding feature, z~p(z) means that the noise obeys the known distribution p(z), and the distribution p(z) can be a Gaussian distribution. D represents the discriminator, and G represents the generator, so the above D(x) and D(G(z)) refer to the scores output by the discriminator for the real image and the synthetic image respectively.

[0114] It can be understood that the better the discriminator is, the closer the above D(x) is usually to 1, and accordingly, the closer D(G(z)) is to 0, so that the discriminator can be trained and optimized by making the above objective function reach the maximum value. Correspondingly, when the generator is better, the closer the synthetic image generated by the generator is to the real image, then D(G(z)) is closer to 1, so that the generator can be optimized by making the above objective function reach the minimum value. That is, in the above training process, the synthetic image generated by the generator needs to be discriminated as a real image by the discriminator as much as possible, and the discriminator needs to recognize the real image and the synthetic image as much as possible, so that the two form adversarial training. During the training process, the parameters of the discriminator and the generator are continuously and dynamically adjusted until both reach a relatively optimal state. The embodiment of the present disclosure can use the final generator as the target synthesizer.

[0115] As another example, consider the diffusion model. This model essentially learns the mapping from Gaussian noise to the true image distribution. This learning process involves two steps: forward diffusion and backward denoising. The diffusion process continuously adds noise to the image until it becomes approximately Gaussian. The denoising process starts with completely random Gaussian noise and gradually removes it to restore the true image. Thus, the trained diffusion model can achieve image synthesis through denoising.

[0116] In the embodiment of the present disclosure, the trained target synthesizer can output a synthesized image based on the input noise. The noise controls the properties of the synthesized image. However, it is unknown what properties of the image are controlled by each dimension (size, distribution) of the noise and it is not interpretable. In other words, the use of the target synthesizer can only achieve unconditional image synthesis and cannot be directly applied to downstream tasks. The user cannot manually edit the input noise to synthesize the desired image, that is, the properties of the synthesized image cannot be controlled. Therefore, the target synthesizer in the embodiment of the present disclosure can be combined with an encoder to achieve control of the image properties through the encoder to achieve conditional image synthesis.

[0117] In the disclosed embodiment, a first image is obtained by inputting preset sample coding features into the synthesizer to be trained, performing image synthesis processing, and adjusting the model parameters of the synthesizer to be trained based on the first image and the first sample image. Training continues based on the adjusted synthesizer to be trained until a second preset stopping condition is reached, at which point the synthesizer to be trained is determined as the target synthesizer. In this way, by training the synthesizer to be trained using the first sample image and adjusting the model parameters, the synthesizer to be trained can learn the distribution of the first sample image and synthesize high-quality images without the need for paired sample data.

[0118] In an optional embodiment, the embodiment of the present disclosure may further include the following steps:

[0119] Step 401: In response to a model generation instruction for a second synthesis task, obtain a third sample image and a second input feature; the second input feature is extracted from the third sample image according to a feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task.

[0120] Among them, the above-mentioned third sample image can also be pre-downloaded from the Internet and can be the same as the above-mentioned second sample image. The above-mentioned second synthesis task refers to a synthesis task with a different feature type from the above-mentioned first synthesis task. The feature type of the second synthesis task and the feature type of the first synthesis task can be set according to actual needs. For example, the feature type of the first synthesis task can be two-dimensional information, and the feature type of the second synthesis task can be three-dimensional information. It can be understood that at this time, the above-mentioned first encoder cannot perform corresponding processing for the input feature of this type of three-dimensional information. Therefore, in the embodiment of the present disclosure, a to-be-trained encoder can be re-trained for the feature type of the second synthesis task.

[0121] Step 402: Based on the target synthesizer, the third sample image, and the second input feature, the encoder to be trained is trained to generate a second encoder that matches the feature type of the second synthesis task; the second encoder is used to generate encoding features of the image based on the second input feature.

[0122] Step 403: Combine the second encoder with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task.

[0123] Among them, the training process for generating the second encoder is the same as the training process of the first encoder. Specifically, the training method for the encoder to be trained can be to use the obtained second input feature as input, and obtain the corresponding encoding feature through the encoder to be trained, so that the encoding feature can be input into the target synthesizer to enable the target synthesizer to generate a synthesized image. Based on the difference between the synthesized image and the corresponding third sample image, the parameters of the encoder to be trained are adjusted and training is continued until the difference between the synthesized image and the third sample image meets the requirements, thereby obtaining the above-mentioned second encoder.

[0124] Figure 4 is a flowchart of another method for generating an image synthesis model according to an exemplary embodiment. Figure 4 As shown, it can be seen that the generation method in the embodiment of the present disclosure may include two processes:

[0125] S1: Training generative models (synthesizers / pre-trained generators) on large amounts of unpaired data.

[0126] S2: Only a small amount of paired data is needed to train the encoder for each task.

[0127] It can be seen that in the embodiment of the present disclosure, for different synthesis tasks, only a small amount of paired data needs to be used to retrain the encoder to realize the application of downstream tasks, and there is no need to collect a large amount of data to train the entire generative network. Therefore, the data cost of model training can be reduced. Since paired data is difficult to obtain and time-consuming to obtain, only a small amount of paired data is needed in the embodiment of the present disclosure. Therefore, the time consumption of obtaining training data is reduced to a certain extent, and the efficiency of obtaining training data is improved, thereby reducing the generation cost of the image synthesis model to a certain extent and improving the generation efficiency.

[0128] In an embodiment of the present disclosure, a third sample image and a second input feature are obtained in response to a model generation instruction for a second synthesis task; the second input feature is extracted from the third sample image according to the feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task; based on the target synthesizer, the third sample image and the second input feature, the encoder to be trained is trained to generate a second encoder that matches the feature type of the second synthesis task; the second encoder is used to generate encoding features of the image based on the second input feature; the second encoder is combined with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task. In this way, for different synthesis tasks, only the encoder needs to be retrained, and there is no need to collect a large amount of data to train the entire generative network, which can reduce the generation cost of the image synthesis model to a certain extent and improve the generation efficiency.

[0129] In an optional embodiment, the embodiment of the present disclosure may further include the following steps:

[0130] Step 501: Based on the feature type of the first synthesis task, select a feature extraction model corresponding to the feature type from preset feature extraction models as a target feature extraction model.

[0131] Step 502: Extract features matching the feature type of the first synthesis task from the second sample image based on the target feature extraction model to obtain the first input features.

[0132] Among them, the above-mentioned feature extraction model corresponds to the feature type of the synthesis task. For example, when the feature type is a two-dimensional segmentation map, the above-mentioned target feature extraction model can be a two-dimensional segmentation model. Correspondingly, when the feature type is a two-dimensional key point, the above-mentioned target feature extraction model can be a key point extraction model.

[0133] Specifically, multiple feature extraction models for different feature types can be pre-set, and identifiers can be set for different models based on different types. In actual applications, the target feature extraction model corresponding to the feature type of the synthesis task can be selected according to the identifier of the feature extraction model. Alternatively, the feature type name can be directly used as the name of the feature extraction model, so that the corresponding target feature extraction model can be selected through the type name and the feature extraction model name.

[0134] Optionally, corresponding input features may be extracted from sample images by manual annotation.

[0135] In the disclosed embodiment, based on the feature type of the first synthesis task, a feature extraction model corresponding to the feature type is selected from preset feature extraction models as the target feature extraction model; and features matching the feature type of the first synthesis task are extracted from the second sample image based on the target feature extraction model to obtain the first input features. In this way, by selecting a feature extraction model that matches the input type of the synthesis task and extracting the corresponding features, the efficiency of obtaining the first input features can be improved, further improving the efficiency of generating the image synthesis model.

[0136] In an optional embodiment, the designated type of image includes a human body image; and the feature type includes at least one of a two-dimensional key point of a human body, a human body segmentation map, and a three-dimensional model of a human body.

[0137] Among them, the specified type of image may include a human body image. Accordingly, the goal of human body image synthesis is to generate a human body image that is difficult to distinguish between true and false, which has many applications in virtual human creation, virtual try-on, content production, etc. According to the different inputs, human body image synthesis can be divided into unconditional image synthesis and conditional image synthesis. The input of unconditional image synthesis is generally noise, and people cannot directly control the properties of the synthesized image (such as: human body posture and movement, etc.). The input of conditional image synthesis is generally a representation of the human body, such as: two-dimensional key points of the human body, three-dimensional shape coefficients of the human body, etc., and its output is a human body image that meets the input characteristics. In the embodiment of the present disclosure, conditional image synthesis can be achieved by combining an encoder and a synthesizer.

[0138] It should be noted that the image types of the first, second and third sample images are the same as the specified types.

[0139] Figure 5 is a schematic diagram showing a feature type according to an exemplary embodiment. Figure 5 As shown, three types of human body key points, human body segmentation map and human body three-dimensional model (3D) are shown. It can be seen that the human body key points and human body segmentation map are different forms of expression of an image, and the human body postures and movements they contain are the same.

[0140] Figure 6 is a schematic diagram showing an image synthesis according to an exemplary embodiment. Figure 6 As shown, the generator refers to the image synthesis model in the embodiment of the present disclosure. It can be seen that the generator synthesizes a synthetic image corresponding to the two-dimensional key points through the input two-dimensional key points, and the posture and other information of the synthetic image are consistent with the posture information represented by the two-dimensional key points.

[0141] In the disclosed embodiment, the designated image type includes a human body image; the feature type includes at least one of a two-dimensional human body key point, a human body segmentation map, and a three-dimensional human body model. Thus, a human body image that meets the requirements can be synthesized by using different human body image feature types.

[0142] Figure 7 is a flow chart of an image synthesis method according to an exemplary embodiment. Figure 7 As shown, the method may include:

[0143] Step 211: Based on the target synthesis task, obtain the target input feature of the target synthesis task and the target feature type corresponding to the target input feature.

[0144] The target synthesis task can be user-created or automatically generated upon receiving a synthesis instruction. Specifically, upon receiving a target synthesis task, the target synthesis task typically includes target input features, i.e., the attribute features of the image to be synthesized by the synthesis task. Furthermore, the corresponding target input type, such as two-dimensional information or three-dimensional information, can be determined based on the target input features.

[0145] Step 212: Based on the target feature type, select an image synthesis model whose encoder matches the target feature type from pre-trained image synthesis models as the target image synthesis model.

[0146] Step 213: Use the target input features as input of the target image synthesis model to obtain a target image synthesized by the target image synthesis model based on the target input features.

[0147] Among them, the above-mentioned pre-trained image synthesis model can be one or more. Specifically, the above-mentioned pre-trained image synthesis model can be trained according to the method provided in the aforementioned image synthesis model generation method embodiment. Furthermore, since the image synthesis model includes an encoder and a target synthesizer, the embodiment of the present disclosure can select an image synthesis model whose encoder matches the target feature type based on the above-mentioned target feature type, so that the target image synthesis model can synthesize the target image based on the target input feature.

[0148] Specifically, after the target input features are input into the target image synthesis model, the encoder in the model will encode the input features into corresponding encoding features, so that the target synthesizer can generate the corresponding target image based on the encoding features.

[0149] Figure 8 is a schematic diagram showing another image synthesis method according to an exemplary embodiment. Figure 8 As shown, the input control conditions 1 to k refer to the above-mentioned input features, which correspond to different feature types. It can be seen that different feature types correspond to different encoders. Furthermore, the encoder can encode different types of input features into noise (that is, the above-mentioned encoding features). It should be noted that the noise output by the encoder is the same as the parameters (distribution type, dimension, etc.) of the sample encoding features used in the synthesizer training process. Furthermore, the pre-trained generator (target synthesizer) synthesizes different target images (synthesized images 1 to k) for the noise generated by different encoders. It can be seen that the same pre-trained generator can be combined with different encoders to obtain image synthesis models for different types.

[0150] Figure 9 is a block diagram of an image synthesis model generation device according to an exemplary embodiment. Figure 9 As shown, the device 60 may include:

[0151] The synthesizer training module 601 is configured to execute, in response to a model generation instruction for a first synthesis task, training a synthesizer to be trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoded features of the image;

[0152] The reference image generation module 602 is configured to generate a reference image corresponding to a first input feature using the target synthesizer and the encoder to be trained; the first input feature is extracted from a second sample image according to a feature type of the first synthesis task; and the number of the second sample images is less than the number of the first sample images.

[0153] A first encoder training module 603 is configured to train the encoder to be trained based on the control image and the second sample image to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate encoding features of the image based on the first input features;

[0154] The first combining module 604 is configured to combine the first encoder with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

[0155] In an optional embodiment, the control image generation module 602 includes:

[0156] a feature input submodule, configured to input the first input feature into the encoder to be trained to obtain an encoding feature of the second sample image corresponding to the first input feature;

[0157] The reference image generation module 602 is specifically configured to input the encoding features of the second sample image into the target synthesizer to generate a reference image;

[0158] The first encoder training module 603 is specifically configured to adjust the model parameters in the encoder to be trained based on the control image and the second sample image, and continue training based on the adjusted encoder to be trained until a first preset stop condition is reached, and the encoder to be trained at the end of training is determined to be the first encoder.

[0159] In an optional embodiment, the synthesizer training module 601 includes:

[0160] a sample feature input submodule configured to input a preset sample coding feature into the synthesizer to be trained, perform image synthesis processing, and obtain a first image;

[0161] The synthesizer parameter adjustment sub-submodule is configured to adjust the model parameters in the synthesizer to be trained based on the first image and the first sample image, and continue training based on the adjusted synthesizer to be trained until a second preset stop condition is reached, and determine the synthesizer to be trained at the end of training as the target synthesizer.

[0162] In an optional embodiment, the device 60 further includes:

[0163] an acquisition module configured to execute, in response to a model generation instruction for a second synthesis task, acquire a third sample image and a second input feature; the second input feature is extracted from the third sample image according to a feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task;

[0164] a second encoder training module configured to train the encoder to be trained based on the target synthesizer, the third sample image, and the second input features to generate a second encoder matching the feature type of the second synthesis task; the second encoder is configured to generate encoding features of an image based on the second input features;

[0165] The second combination module is configured to combine the second encoder with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task.

[0166] In an optional embodiment, the device 60 further includes:

[0167] A selection module is configured to execute, based on the feature type of the first synthesis task, selecting a feature extraction model corresponding to the feature type from preset feature extraction models as a target feature extraction model;

[0168] The feature extraction module is configured to extract features matching the feature type of the first synthesis task from the second sample image based on the target feature extraction model to obtain the first input feature.

[0169] In an optional embodiment, the designated type of image includes a human body image; and the feature type includes at least one of a two-dimensional key point of a human body, a human body segmentation map, and a three-dimensional model of a human body.

[0170] In summary, the image synthesis model generation device provided by the embodiment of the present disclosure trains the synthesizer to be trained based on the first sample image in response to the model generation instruction for the first synthesis task to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoding features of the image; a control image corresponding to the first input feature is generated through the target synthesizer and the encoder to be trained; the first input feature is extracted from the second sample image according to the feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images; based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate the encoding features of the image based on the first input feature; the first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task. In this way, since the synthesizer is independent of the input type of the synthesis task, the synthesizer only needs to synthesize the image based on the encoding features of the image. Therefore, there is no need to use a large amount of paired data to train the synthesizer to be trained according to the synthesis task. It is only necessary to obtain a small amount of second sample images and first input features to train the encoder to be trained. To a certain extent, it can reduce the amount of paired sample data required in the process of generating the image synthesis model and reduce the generation cost of the model.

[0171] At the same time, by adaptively training the synthesizer and encoder separately, the synthesis process and the synthesis task are decoupled. When a new synthesis task is received, there is no need to retrain the synthesizer, only the encoder needs to be trained, thus achieving the reusability of the synthesizer for different synthesis tasks.

[0172] Figure 10 is a block diagram of an image synthesis device according to an exemplary embodiment. Figure 10 As shown, the device 70 may include:

[0173] The target acquisition module 701 is configured to execute a target synthesis task, obtain a target input feature of the target synthesis task and a target feature type corresponding to the target input feature;

[0174] A selection module 702 is configured to select, based on the target feature type, an image synthesis model whose encoder matches the target feature type from pre-trained image synthesis models as a target image synthesis model;

[0175] The image acquisition module 703 is configured to use the target input features as input to the target image synthesis model, and obtain a target image synthesized by the target image synthesis model based on the target input features.

[0176] According to one embodiment of the present disclosure, an electronic device is provided, comprising: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to implement the steps in the image synthesis model generation method or image synthesis method in any of the above embodiments when executed.

[0177] According to one embodiment of the present disclosure, a computer-readable storage medium is also provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the steps in the image synthesis model generation method or image synthesis method in any of the above embodiments.

[0178] According to one embodiment of the present disclosure, a computer program product is also provided, which includes readable program instructions. When the readable program instructions are executed by a processor of an electronic device, the electronic device can execute the steps in the image synthesis model generation method or image synthesis method in any of the above embodiments.

[0179] Figure 11 8 is a block diagram of an apparatus for generating an image synthesis model according to an exemplary embodiment. The apparatus 800 may include a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output interface 812, a sensor component 814, a communication component 816, and a processor 820. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned image synthesis model generation method. In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 804 including instructions, and the above-mentioned instructions can be executed by the processor 820 of the apparatus 800 to complete the above-mentioned method. Optionally, the storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0180] Figure 12 This is a block diagram illustrating another apparatus for generating an image synthesis model according to an exemplary embodiment. Apparatus 900 may include a processing component 922, a memory 932, an input / output interface 958, a network interface 950, and a power supply component 926. Apparatus 900 may be provided as a server. The application stored in memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, processing component 922 is configured to execute the instructions to perform the aforementioned method for generating an image synthesis model.

[0181] The user information (including but not limited to the user's device information, user personal information, etc.) and related data involved in this disclosure are all information authorized by the user or authorized by all parties.

[0182] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0183] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating an image synthesis model, characterized in that: The method comprises: In response to a model generation instruction for a first synthesis task, training a synthesizer to be trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoding features of the image; Generate, by means of the target synthesizer and the encoder to be trained, a control image corresponding to a first input feature; the first input feature is extracted from a second sample image according to a feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images; the designated type of image includes a human body image; the feature type includes at least one of a two-dimensional human body key point, a human body segmentation map, and a three-dimensional human body model; Based on the control image and the second sample image, the encoder to be trained is trained to generate a first encoder that matches the feature type of the first synthesis task; the first encoder is used to generate encoding features of the image based on the first input features; The first encoder is combined with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

2. The method according to claim 1, characterized in that The step of generating a control image corresponding to the first input feature by using the target synthesizer and the encoder to be trained includes: Inputting the first input feature into the encoder to be trained to obtain an encoding feature of the second sample image corresponding to the first input feature; inputting the encoded features of the second sample image into the target synthesizer to generate a control image; The training of the encoder to be trained based on the control image and the second sample image includes: adjusting the model parameters in the encoder to be trained based on the control image and the second sample image, and continuing training based on the adjusted encoder to be trained until a first preset stop condition is reached, and determining the encoder to be trained at the end of training as the first encoder.

3. The method according to claim 1, characterized in that The step of training the synthesizer to be trained based on the first sample image to generate a target synthesizer includes: Inputting the preset sample coding features into the synthesizer to be trained, performing image synthesis processing, and obtaining a first image; Based on the first image and the first sample image, the model parameters in the synthesizer to be trained are adjusted, and training is continued based on the adjusted synthesizer to be trained until a second preset stop condition is reached, and the synthesizer to be trained at the end of training is determined as the target synthesizer.

4. The method according to claim 1, wherein The method further comprises: In response to a model generation instruction for a second synthesis task, obtaining a third sample image and a second input feature; the second input feature is extracted from the third sample image according to a feature type of the second synthesis task; the feature type of the second synthesis task is different from the feature type of the first synthesis task; Based on the target synthesizer, the third sample image, and the second input feature, the encoder to be trained is trained to generate a second encoder that matches the feature type of the second synthesis task; the second encoder is used to generate encoding features of the image based on the second input feature; The second encoder is combined with the target synthesizer to obtain a second image synthesis model corresponding to the feature type of the second synthesis task.

5. The method according to claim 1, wherein The method further comprises: Based on the feature type of the first synthesis task, selecting a feature extraction model corresponding to the feature type from preset feature extraction models as a target feature extraction model; Features matching the feature type of the first synthesis task are extracted from the second sample image based on the target feature extraction model to obtain the first input features.

6. An image synthesis model generation device, characterized in that: The device comprises: a synthesizer training module configured to execute, in response to a model generation instruction for a first synthesis task, training a synthesizer to be trained based on a first sample image to generate a target synthesizer; the target synthesizer is used to synthesize a specified type of image based on the encoded features of the image; a reference image generation module configured to generate, through the target synthesizer and the encoder to be trained, a reference image corresponding to a first input feature; the first input feature is extracted from a second sample image according to a feature type of the first synthesis task; the number of the second sample images is less than the number of the first sample images; the designated type of image includes a human body image; the feature type includes at least one of a two-dimensional human body key point, a human body segmentation map, and a three-dimensional human body model; a first encoder training module configured to train the encoder to be trained based on the control image and the second sample image to generate a first encoder matching the feature type of the first synthesis task; the first encoder is configured to generate encoding features of an image based on the first input features; A first combination module is configured to combine the first encoder with the target synthesizer to obtain a first image synthesis model corresponding to the feature type of the first synthesis task.

7. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Semantic image synthesis for generating substantially photorealistic images using neural networks

    CN111489412A

  • Speech synthesis method and speech synthesis system

    CN112908294A