Method and apparatus for generating image, method and apparatus for constructing stylization model, and device

By using the diffusion model to generate diverse stylized images and training the generative adversarial network, the problems of high cost and poor training of existing stylized models are solved, and high-quality and high-reliability stylized image generation is achieved.

WO2025130788A1PCT designated stage expired Publication Date: 2025-06-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139366
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The training cost of existing stylized models is high, poor in effect, and poor in reliability. The main reason is that the number of style graphs is small, which leads to the difficulty of data collection and the reliability of model training.

Method used

By acquiring the target diffusion model, the diffusion model is used to generate diverse stylized images and pair them with the original image structure to train the generative adversarial network to generate high-quality stylized images.

Benefits of technology

It realizes the reduction of the cost of acquisition of stylized images, improves the quantity and quality of stylized images, and enhances the reliability and stability of stylized models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139366_26062025_PF_FP_ABST
    Figure CN2024139366_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for generating an image, a method and apparatus for constructing a stylization model, and a device. The method for generating an image comprises: acquiring a target diffusion model; on the basis of an acquired first original image, using the target diffusion model to generate a first stylized image corresponding to the first original image, so as to construct first pairing data on the basis of the first original image and the first stylized image; using the first pairing data to train an initial first generative adversarial network, so as to obtain a target generative adversarial network; and on the basis of an acquired target original image, using the target generative adversarial network to generate a target stylized image corresponding to the target original image.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, stylized model construction method, device and equipment

[0001] This application claims priority to the Chinese invention patent application entitled “Image generation method, method, device and equipment for constructing stylized model” and filed on December 22, 2023, with application number 202311791067.7. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of computer technology, and in particular to an image generation method, a stylized model construction method, an apparatus, and a device. Background Art

[0003] In order to meet the personalized needs of users and enrich the visual effects of images, stylized models are used in more and more image processing scenarios to stylize the user's input images, thereby obtaining the stylized images required by the user. Summary of the Invention

[0004] The present disclosure provides an image generation method, a stylized model construction method, an apparatus, and a device.

[0005] An embodiment of the present disclosure provides an image generation method, including: obtaining a target diffusion model; based on an obtained first original image, generating a first stylized image corresponding to the first original image using the target diffusion model, so as to construct first paired data based on the first original image and the first stylized image; training an initial first generative adversarial network using the first paired data to obtain a target generative adversarial network; and based on the obtained target original image, generating a target stylized image corresponding to the target original image using the target generative adversarial network.

[0006] An embodiment of the present disclosure provides a method for constructing a stylized model, comprising: obtaining a training dataset; wherein the training dataset includes a target original image and a target stylized image corresponding to the target original image, and the target stylized image is obtained using the aforementioned image generation method; using the training dataset to train a preset initial network model to construct a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model.

[0007] The presently disclosed embodiment also provides an image generation device, comprising: a model acquisition module for acquiring a target diffusion model; a first generation module for generating, based on an acquired first original image, a first stylized image corresponding to the first original image using the target diffusion model, so as to construct first paired data based on the first original image and the first stylized image; a network training module for training an initial first generative adversarial network using the first paired data to obtain a target generative adversarial network; and a second generation module for generating, based on the acquired target original image, a target stylized image corresponding to the target original image using the target generative adversarial network.

[0008] The disclosed embodiments also provide a device for constructing a stylized model, comprising: a data set acquisition module for acquiring a training data set; wherein the training data set includes a target original image and a target stylized image corresponding to the target original image, and the target stylized image is obtained using the aforementioned image generation method; a model construction module for training a preset initial network model using the training data set to construct a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model.

[0009] An embodiment of the present disclosure also provides an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the image generation method or stylized model construction method provided in the embodiment of the present disclosure.

[0010] The embodiments of the present disclosure further provide a computer-readable storage medium, which stores a computer program for executing the image generation method or the stylized model construction method provided in the embodiments of the present disclosure.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0013] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] FIG1 is a schematic diagram of a flow chart of an image generation method provided by an embodiment of the present disclosure;

[0015] FIG2 is a schematic diagram of a flow chart of a method for constructing a stylized model provided by an embodiment of the present disclosure;

[0016] FIG3 is a schematic diagram of a system for generating a stylized image according to an embodiment of the present disclosure;

[0017] FIG4 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure;

[0018] FIG5 is a schematic structural diagram of a stylized model construction device provided by an embodiment of the present disclosure;

[0019] FIG6 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The inventors have found through research that the training cost of existing stylized models is high, the trained models are not effective, and the reliability is poor; the main reason is that the number of existing style maps is usually small, so it is difficult to collect model training data (including paired original images and corresponding style maps), and the reliability of the models trained with a small amount of training data is usually poor. In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image generation method, a stylized model construction method, an apparatus and a device, and the scheme of the present disclosure will be further described below. It should be noted that, in the absence of conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0022] Figure 1 is a flow chart of an image generation method provided by an embodiment of the present disclosure. The method can be executed by an image generation device, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. As shown in Figure 1, the method mainly includes the following steps S102 to S108:

[0023] Step S102: Acquire a target diffusion model.

[0024] In practical applications, a preset diffusion model can be directly used to obtain the target diffusion model. This preset diffusion model can be a model from an open source library or a previously pre-trained model; there is no restriction on the source of the preset diffusion model. To further enhance the diversity of the diffusion model output images, multiple preset diffusion models can be fused, and the resulting new diffusion model can be used as the target diffusion model.

[0025] Step S104 : Based on the acquired first original image, a first stylized image corresponding to the first original image is generated using a target diffusion model, so as to construct first pairing data based on the first original image and the first stylized image.

[0026] The first original image can serve as the initial input image for the target diffusion model, allowing the target diffusion model to perform stylization processing on the first original image to obtain a first stylized image. In practical applications, the first original image can also be updated. Based on the updated first original image, the target diffusion model can generate multiple first stylized images, thereby increasing the diversity of the first stylized images. For ease of implementation, the first stylized image output by the target diffusion model can also be used to update the first original image. For example, the first stylized image output by the target diffusion model in a previous operation can be used as the first original image input in a subsequent operation, thereby enhancing the degree of stylization. Assuming that the target diffusion model performs multiple image generation operations, when constructing the first paired data based on the first original image and the first stylized image, the model input image and model output image corresponding to each image generation operation can be combined to obtain the first paired data. Furthermore, the model input image and model output image corresponding to different image generation operations can be combined as needed to obtain the first paired data. The specific application is flexible and not limited here.

[0027] Step S106: Use the first paired data to train an initial first generative adversarial network to obtain a target generative adversarial network.

[0028] Through the above-mentioned method, a large amount of diverse first paired data can be obtained, and the effect of using the first paired data to train the initial first generative adversarial network is better and more robust. And because the images generated by the generative adversarial network usually have characteristics such as strong stability, low bad example rate, and reduced randomness, the stylized images generated by the trained first generative adversarial network have a low bad example rate and high quality, and the overall stylized effect is more stable and controllable. Then, the embodiment of the present disclosure can further obtain the required target generative adversarial network based on the obtained trained first generative adversarial network. In some embodiments, the first generative adversarial network can be directly used as the target generative adversarial network. In other embodiments, the images generated by the first generative adversarial network can also be used to train a lightweight generative adversarial network to obtain the target generative adversarial network. In this way, it can be ensured that the target generative adversarial network can be applied to devices such as mobile terminals, shortening the processing time and resource consumption of the target generative adversarial network, and achieving the effect of real-time stylized processing as much as possible.

[0029] In step S108, based on the acquired target original image, a target generative adversarial network is used to generate a target stylized image corresponding to the target original image. The target original image may be the same as or different from the first original image. The target stylized image may be provided directly to the user. For example, the user may directly use the target generative adversarial network to generate the desired target stylized image. In this case, the target original image may be an image uploaded by the user. The target stylized image and the target original image may also be used to form paired data to train other models. The disclosed embodiments do not limit the use of the target stylized image.

[0030] The above method provided by the embodiment of the present disclosure fully takes into account that the images generated by the diffusion model usually have characteristics such as strong diversity and strong randomness, while the images generated by the generative adversarial network usually have characteristics such as strong stability and low bad example rate. Therefore, the embodiment of the present disclosure can first generate a variety of first stylized images with the help of the image generation ability of the diffusion model, thereby better ensuring the quantity and diversity of the first paired data. The first paired data is then used to train the first generative adversarial network, and the obtained target generative adversarial network is then used to stably and controllably generate stylized images that meet the expectations, thereby effectively ensuring the quality of the target stylized images finally obtained. Using the above image generation method, there is no need to spend time and energy to collect stylized images. Instead, a large number of required stylized images can be directly generated, ensuring the number of stylized images and reducing the cost of obtaining stylized images.

[0031] In some implementations, the above step S102, i.e., the step of obtaining the target diffusion model, may be performed with reference to the following steps 1 and 2:

[0032] Step 1: Obtain multiple preset diffusion models. As mentioned above, these models can be from open-source libraries or pre-trained models. Different preset diffusion models can have the same or different structures. Preset diffusion models with the same structure may have different network parameters, and the training datasets used to train the preset diffusion models may also differ. Therefore, different preset diffusion models can achieve different stylized effects.

[0033] Step 2: Fusion processing is performed based on multiple preset diffusion models to obtain a target diffusion model. By fusing multiple preset diffusion models, the stylized effect achievable by the target diffusion model can be a comprehensive presentation of the stylized effects achievable by multiple preset diffusion models, further enhancing the diversity of stylized effects.

[0034] For ease of processing, in some specific implementation examples, different preset diffusion models have the same network structure but different network parameters. Based on this, when implementing step 2, the network parameters of multiple preset diffusion models can be fused to obtain a target diffusion model. The network structure of the target diffusion model is the same as that of the preset diffusion models. For example, assuming there are three preset diffusion models, namely Model A, Model B, and Model C, the fusion process can be implemented by determining the parameters of the first network in the target diffusion model based on the parameters of the first network in Model A, the parameters of the second network in the target diffusion model based on the parameters of the second network in Model B, and the parameters of the third network in the target diffusion model based on the parameters of the third network in Model C. Alternatively, the fusion process can be implemented by determining the network parameters of the target diffusion model using a weighted average of the parameters of Models A, B, and C, or other fusion methods. The above descriptions are illustrative and not limiting. By fusing the parameters of the preset diffusion models, the resulting target diffusion model can efficiently and conveniently integrate the stylized effects of multiple preset diffusion models.

[0035] In some embodiments, the step of generating a first stylized image corresponding to the first original image using the target diffusion model based on the acquired first original image in step S104 may be performed with reference to the following steps a to c:

[0036] Step a, obtaining the content information of the first original image and the original prompt information corresponding to the first original image. The content information may be, for example, the category information of the content contained in the first original image detected by the algorithm. The category granularity can be set according to demand, such as major categories such as person categories and animal categories, or the person categories can be divided into subcategories such as age categories and gender categories, and the animal categories can be divided into subcategories such as cats and dogs. When the target diffusion model generates the first stylized image corresponding to the first original image, prompt information is usually required, and the prompt information is used to indicate the relevant information of the stylized image that the target diffusion model needs to generate. The more detailed and specific the information described in the prompt information, the more the stylized image generated by the target diffusion model can meet the demand. The way the target diffusion model generates the corresponding stylized image based on the prompt information can refer to the relevant technology and will not be repeated here.

[0037] Step b: Expand the original prompt information based on the content information to obtain target prompt information. For example, the content information can be added to the original prompt information according to a preset format to obtain the target prompt information. This method allows the target prompt information to include the content information of the first original image, thereby enhancing the correlation between the first stylized image output by the target diffusion model and the first original image.

[0038] Exemplarily, step b may be performed with reference to the following steps b1 to b3:

[0039] Step b1 determines the content attributes of the first original image based on the content information of the first original image. In practical applications, an attribute detection algorithm can be used to determine the content attributes of the first original image. Attributes can be simply understood as the category to which the image content belongs. Taking a person as an example, person attributes may include gender, age, posture, and decoration. A variety of attributes can be flexibly defined based on specific needs and are not limited here.

[0040] Step b2: Obtain the description text corresponding to the content attribute. In practical applications, the description text corresponding to each content attribute can be pre-set. The description text can be presented in a preset format that matches the text format in the prompt information for easier model processing.

[0041] Step b3: Expand the original prompt information based on the description text to obtain target prompt information. The description text can be combined with the original prompt information to obtain the target prompt information.

[0042] In step c, a target diffusion model is used to generate a first stylized image corresponding to the first original image based on the first original image and the target hint information. Both the first original image and the target hint information serve as inputs to the target diffusion model. Because the target hint information includes the content attribute information of the first original image, the first stylized image output by the target diffusion model is more closely related to the first original image and more closely meets user needs.

[0043] In some embodiments, the step S104 of generating a first stylized image corresponding to the first original image using the target diffusion model based on the acquired first original image may also be performed by referring to the following steps: performing multiple image generation operations using the target diffusion model based on the acquired first original image to obtain a first stylized image corresponding to each image generation operation; wherein the input image corresponding to the first image generation operation performed by the target diffusion model is the first original image, and the input image corresponding to the non-first image generation operation performed by the target diffusion model is the first stylized image obtained by the previous image generation operation. In this manner, the target diffusion model can be used to generate stylized images multiple times continuously, thereby efficiently and reliably enhancing and controlling the degree of stylization. Furthermore, in practical applications, the enhancement degree parameters corresponding to different image generation operations can be the same or different, and can be flexibly set according to needs. By adjusting the enhancement degree parameters, the stylization degree of the output image of the target diffusion model can be further controlled, resulting in a variety of first stylized images.

[0044] In some embodiments, step S106, i.e., the step of training the initial first generative adversarial network using the first paired data to obtain the target generative adversarial network, can be performed with reference to steps A to C below:

[0045] In step A, an initial first generative adversarial network is trained using the first paired data to obtain a trained first generative adversarial network. Because a large amount of diverse first paired data can be obtained through the aforementioned method, the trained first generative adversarial network obtained based on this data is more robust.

[0046] In step B, based on a preset second original image, second paired data is obtained using the trained first generative adversarial network. Compared to the diffusion model, the generative adversarial network generates images with moderately reduced randomness, greater stability, and a lower failure rate. To improve the stability of the stylized effect and reduce the failure rate, the disclosed embodiments can further generate stylized images using the trained first generative adversarial network, thereby achieving data distillation and obtaining second paired data of higher quality than the first paired data.

[0047] In some specific implementation examples, a trained first generative adversarial network can perform multiple image generation operations based on a preset second original image to obtain a second stylized image corresponding to each image generation operation, and second paired data can be constructed based on the second original image and the second stylized image. The input image corresponding to the first image generation operation performed by the trained first generative adversarial network is the second original image, and the input image corresponding to the non-first image generation operation performed by the trained first generative adversarial network is the second stylized image obtained by the previous image generation operation. By having the first generative adversarial network perform multiple image generation operations, the quality of the second stylized image can be effectively guaranteed while further improving the stylization level of the resulting second stylized image.

[0048] When performing multiple image generation operations using the trained first generative adversarial network, the following steps can be used: for each image generation operation, obtain the input image corresponding to the image generation operation performed by the trained first generative adversarial network; determine a target region within the input image corresponding to the image generation operation; and perform stylized processing on the target region using the trained first generative adversarial network to obtain a second stylized image corresponding to the image generation operation. Specifically, different image generation operations correspond to different target regions, such as the face region, background region, or body region. The regions can be flexibly divided according to needs. For example, for input image B, multiple segmentation maps B1, B2, ..., Bn can be generated based on input image B. Each image generation operation processes a different segmentation map. For example, the first image generation operation primarily performs stylized processing on segmentation map 1 within input image B, while the second image generation operation primarily performs stylized processing on segmentation map 2 within input image B. Similarly, through the above method, a second stylized image with stylized local regions can be obtained, greatly enhancing the flexibility of the stylized processing and the diversity of the stylized images.

[0049] In step C, the initial second generative adversarial network is trained using the second paired data to obtain a target generative adversarial network based on the trained second generative adversarial network. In practical applications, the trained second generative adversarial network can be directly used as the target generative adversarial network, or other network layers can be added as needed, which is not limited here.

[0050] The second GAN has fewer network parameters than the first. This makes it a lightweight network, requiring fewer network parameters and consuming fewer computing resources. This significantly increases image processing speed and reduces device hardware requirements, allowing it to run flexibly on mobile devices and achieve real-time stylization.

[0051] Furthermore, the second paired data obtained through this method is high-quality and produces a stable stylized effect. Because it is generated through model generation, it can effectively meet the quantity requirements. The second generative adversarial network trained using this method is highly reliable and can output stylized images that meet expectations.

[0052] In practical applications, the aforementioned generative adversarial networks (such as the first generative adversarial network and the second generative adversarial network) may include an encoder, a decoder, and a discriminator. During training, random sampling can be performed from the corresponding paired data, and the original images in the paired data can be input into the encoder and decoder in turn to gradually complete the style conversion process, thereby outputting a stylized image; the structural loss can be calculated for the stylized images in the paired data and the stylized images output by the model, and the structural loss may include L1 loss and Lpips loss (image perception loss), etc., and then the stylized images in the paired data and the stylized images output by the model can be input into the discriminator to calculate the adversarial loss, and the total loss is calculated based on the above structural loss and adversarial loss. The network parameters of the generative adversarial network are adjusted according to the total loss, and the network training is performed with the goal of reducing the total loss, and finally a generative adversarial network that meets the requirements is obtained. In addition, the structures of the first generative adversarial network and the second generative adversarial network are different, and the training parameters can also be adaptively adjusted according to their respective situations, which will not be repeated here.

[0053] In summary, the image generation method provided by the embodiments of the present disclosure fully takes into account that images generated by the diffusion model generally have characteristics such as strong diversity and strong randomness, while images generated by the generative adversarial network generally have characteristics such as strong stability and low bad example rate. Therefore, the embodiments of the present disclosure can first generate a variety of first stylized images with the help of the image generation capability of the diffusion model, thereby better ensuring the quantity and diversity of the first paired data, and then use the first paired data to train the first generative adversarial network, and then use the obtained target generative adversarial network to stably generate stylized images that meet the expectations, thereby effectively ensuring the quality and quantity of the target stylized images finally obtained.

[0054] The technical solutions provided by the embodiments of the present disclosure fully consider that images generated by diffusion models typically have characteristics such as high diversity and randomness, while images generated by generative adversarial networks typically have characteristics such as strong stability and a low rate of bad examples. Therefore, the embodiments of the present disclosure can first generate a variety of first stylized images by leveraging the image generation capabilities of the diffusion model, thereby effectively ensuring the quantity and diversity of the first paired data. The first paired data is then used to train the first generative adversarial network, and the resulting target generative adversarial network is then used to stably generate desired stylized images, thereby effectively ensuring the quality of the final target stylized image. Furthermore, the above image generation method can efficiently acquire a large number of stylized images, thereby reducing the cost of acquiring training datasets. The stylized model obtained by training the initial network model using a large number of high-quality training datasets is highly reliable.

[0055] Based on the above, the present disclosure also provides a method for constructing a stylized model. Figure 2 is a flow chart of a method for constructing a stylized model provided by the present disclosure. The method can be performed by a stylized model construction device, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. As shown in Figure 2, the method mainly includes the following steps S202 to S204:

[0056] Step S202: Obtain a training dataset; wherein the training dataset includes a target original image and a target stylized image corresponding to the target original image, and the target stylized image is obtained using any of the aforementioned image generation methods. For details, please refer to the aforementioned related content and will not be repeated here. Through the above method, a high-quality training dataset can be obtained efficiently and reliably, and the amount of data in the training dataset can be effectively guaranteed, which can greatly reduce training costs.

[0057] In step S204, the preset initial network model is trained using the training data set to construct a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model. The disclosed embodiment does not limit the structure of the initial network model, and any network model that requires the use of stylized images as training data for training can be used. In actual applications, the trained initial network model can be directly used as the stylized model, or other network layers can be added as needed based on this basis, which is not limited here.

[0058] The above method provided by the embodiment of the present disclosure can further obtain a training data set containing the target original image and the corresponding target stylized image based on the aforementioned image generation method. The stylized model obtained by training the initial network model using a large number of high-quality training data sets has strong reliability.

[0059] The stylization model can run on the server or the user side, and can perform stylization processing on the user's input image, thereby presenting a stylized image with the target style to meet user needs.

[0060] Referring to FIG3 , a schematic diagram of a stylized image generation system is shown, which primarily includes a diffusion model selection unit, a model adjustment unit, a model fusion unit, a paired data generation unit, and a model training unit. The aforementioned units can be used to obtain a desired stylized model, which can be used to stylize a user's input image, thereby obtaining the desired stylized image. The diffusion model selection unit can select a desired diffusion model from preset diffusion models, the model adjustment unit can adjust the model's input parameters (such as prompt information, stylization level parameters, etc.), and the model fusion unit can perform parameter fusion on multiple preset diffusion models. In practical applications, the model adjustment unit and / or model fusion unit can be flexibly applied to process the selected preset diffusion model as needed to obtain a target diffusion model. Based on this, the paired data generation unit can generate paired data using the target diffusion model, and the model training unit can train an initial network model, such as a generative adversarial network, based on the paired data, thereby obtaining a desired stylized model. In practical applications, the stylized model can directly perform style conversion on the user input image to obtain a stylized image. In addition, the stylized image generated by the stylized model can be used to generate paired data for training other models, which is not limited here.

[0061] In practical applications, the stylized image generation system may further include, for example, a pre-processing unit (e.g., capable of pre-processing user input images) and / or a post-processing unit (e.g., capable of post-processing the stylized image output by the model), without limitation herein. The different units included in the stylized image generation system may be provided on different devices. For example, the diffusion model selection unit, model adjustment unit, model fusion unit, paired data generation unit, and model training unit may be provided on the server side, while the ultimately generated stylized model or pre-processing unit and post-processing unit may be provided on the user side. The specific configuration is flexible and without limitation herein.

[0062] The above technical solution provided by the embodiment of the present disclosure can obtain a large number of high-quality stylized images as pairing data, thereby training a more reliable stylized model, and then using the stylized model to obtain stylized images with better effects and better meet user needs.

[0063] Corresponding to the aforementioned image generation method, FIG4 is a schematic structural diagram of an image generation device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can generally be integrated into an electronic device. As shown in FIG4 , the image generation device includes:

[0064] Model acquisition module 402, used to acquire a target diffusion model;

[0065] A first generating module 404 is configured to generate a first stylized image corresponding to the first original image using a target diffusion model based on the acquired first original image, so as to construct first paired data based on the first original image and the first stylized image;

[0066] A network training module 406 is configured to train an initial first generative adversarial network using the first paired data to obtain a target generative adversarial network;

[0067] The second generation module 408 is configured to generate a target stylized image corresponding to the target original image by using a target generative adversarial network based on the acquired target original image.

[0068] The above-mentioned device provided by the embodiment of the present disclosure fully considers that images generated by the diffusion model generally have characteristics such as strong diversity and strong randomness, while images generated by the generative adversarial network generally have characteristics such as strong stability and low bad example rate. Therefore, the embodiment of the present disclosure can first generate a variety of first stylized images with the image generation capability of the diffusion model, thereby better ensuring the quantity and diversity of the first paired data, and then use the first paired data to train the first generative adversarial network, and then use the obtained target generative adversarial network to stably generate stylized images that meet the expectations, thereby effectively ensuring the quality of the target stylized images finally obtained. Using the above-mentioned image generation method, there is no need to spend time and energy to collect stylized images. Instead, a large number of required stylized images can be directly generated, ensuring the number of stylized images and reducing the cost of obtaining stylized images.

[0069] In some implementations, the model acquisition module 402 is specifically configured to: acquire a plurality of preset diffusion models; and perform fusion processing based on the plurality of preset diffusion models to obtain a target diffusion model.

[0070] In some embodiments, different preset diffusion models have the same network structure but different network parameters; the model acquisition module 402 is specifically used to: fuse the network parameters of the multiple preset diffusion models to obtain a target diffusion model; wherein the network structure of the target diffusion model is the same as the network structure of the preset diffusion model.

[0071] In some embodiments, the first generation module 404 is specifically used to: obtain content information of the obtained first original image and original prompt information corresponding to the first original image; expand the original prompt information based on the content information to obtain target prompt information; and generate a first stylized image corresponding to the first original image using the target diffusion model based on the first original image and the target prompt information.

[0072] In some embodiments, the first generation module 404 is specifically used to: determine the content attributes of the first original image based on the content information of the first original image; obtain the description text corresponding to the content attributes; and expand the original prompt information based on the description text to obtain target prompt information.

[0073] In some embodiments, the first generation module 404 is specifically configured to: perform multiple image generation operations using the target diffusion model based on the acquired first original image to obtain a first stylized image corresponding to each image generation operation; wherein the input image corresponding to the first image generation operation performed by the target diffusion model is the first original image, and the input image corresponding to the non-first image generation operation performed by the target diffusion model is the first stylized image obtained by the previous image generation operation.

[0074] In some embodiments, the network training module 406 is specifically used to: use the first pairing data to train an initial first generative adversarial network to obtain a trained first generative adversarial network; based on a preset second original image, obtain second pairing data through the trained first generative adversarial network; use the second pairing data to train the initial second generative adversarial network to obtain a target generative adversarial network based on the trained second generative adversarial network; wherein the network parameters of the second generative adversarial network are less than the network parameters of the first generative adversarial network.

[0075] In some embodiments, the network training module 406 is specifically used to: based on a preset second original image, perform multiple image generation operations through the trained first generative adversarial network to obtain a second stylized image corresponding to each image generation operation, and construct second pairing data based on the second original image and the second stylized image; wherein, the input image corresponding to the first image generation operation performed by the trained first generative adversarial network is the second original image, and the input image corresponding to the non-first image generation operation performed by the trained first generative adversarial network is the second stylized image obtained by the previous image generation operation.

[0076] In some embodiments, the network training module 406 is specifically used to: for each image generation operation, obtain the input image corresponding to the image generation operation performed by the trained first generative adversarial network; determine the target area in the input image corresponding to the image generation operation; and stylize the target area through the trained first generative adversarial network to obtain a second stylized image corresponding to the image generation operation.

[0077] In some embodiments, different image generation operations correspond to different target areas.

[0078] The image generation device provided in the embodiments of the present disclosure can execute the image generation method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0079] Corresponding to the aforementioned image generation method, FIG5 is a schematic structural diagram of a stylized model construction apparatus provided by an embodiment of the present disclosure. The apparatus may be implemented by software and / or hardware and may generally be integrated into an electronic device. As shown in FIG5 , the stylized model construction apparatus includes:

[0080] The data set acquisition module 502 is used to acquire a training data set; wherein the training data set includes a target original image and a target stylized image corresponding to the target original image, and the training data set is obtained using the aforementioned image generation method;

[0081] The model construction module 504 is used to train a preset initial network model using a training data set to construct a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model.

[0082] The above-mentioned device provided by the embodiment of the present disclosure can further obtain a training data set including the target original image and the corresponding target stylized image based on the aforementioned image generation method. The stylized model obtained by training the initial network model using a large number of high-quality training data sets has strong reliability.

[0083] The stylized model construction device provided in the embodiments of the present disclosure can execute the stylized model construction method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0084] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device embodiment can refer to the corresponding process in the method embodiment, and will not be repeated here.

[0085] An embodiment of the present disclosure provides an electronic device, which includes: a storage device storing a computer program; and a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.

[0086] Reference is now made to FIG6 , which illustrates a schematic diagram of the structure of an electronic device 600 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG6 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.

[0087] As shown in Figure 6, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0088] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0089] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0090] In addition to the above-mentioned methods and devices, the embodiments of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the image processing method provided by the embodiments of the present disclosure. The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present disclosure, the programming languages ​​including object-oriented programming languages ​​such as Java, C++, etc., and also conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0091] In addition, the embodiments of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the image generation method or stylized model construction method provided by the embodiments of the present disclosure.

[0092] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0093] The embodiments of the present disclosure further provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the image generation method or the stylized model construction method in the embodiments of the present disclosure.

[0094] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0095] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0096] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0097] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0098] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0099] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating an image, comprising: Obtain target diffusion model; Based on the acquired first original image, generating a first stylized image corresponding to the first original image by using the target diffusion model, so as to construct first pairing data based on the first original image and the first stylized image; Using the first paired data to train an initial first generative adversarial network to obtain a target generative adversarial network; Based on the acquired target original image, the target generative adversarial network is used to generate a target stylized image corresponding to the target original image.

2. The method according to claim 1, wherein: The obtaining of the target diffusion model comprises: Get multiple preset diffusion models; A fusion process is performed based on the multiple preset diffusion models to obtain a target diffusion model.

3. The method according to claim 2, wherein: The different preset diffusion models have the same network structure but different network parameters; and the fusion processing based on the multiple preset diffusion models to obtain the target diffusion model includes: The network parameters of the plurality of preset diffusion models are fused to obtain a target diffusion model; wherein the network structure of the target diffusion model is the same as the network structure of the preset diffusion model.

4. The method according to claim 1, wherein: The step of generating a first stylized image corresponding to the first original image by using the target diffusion model based on the acquired first original image includes: Acquire content information of the acquired first original image and original prompt information corresponding to the first original image; Expanding the original prompt information based on the content information to obtain target prompt information; Based on the first original image and the target prompt information, a first stylized image corresponding to the first original image is generated using the target diffusion model.

5. The method according to claim 4, wherein: The expanding the original prompt information based on the content information to obtain target prompt information includes: determining a content attribute of the first original image based on content information of the first original image; Obtaining description text corresponding to the content attribute; The original prompt information is expanded based on the description text to obtain target prompt information.

6. The method according to claim 1, wherein: The step of generating a first stylized image corresponding to the first original image by using the target diffusion model based on the acquired first original image includes: Based on the acquired first original image, performing multiple image generation operations using the target diffusion model to obtain a first stylized image corresponding to each image generation operation; The input image corresponding to the first image generation operation performed by the target diffusion model is the first original image, and the input image corresponding to the non-first image generation operation performed by the target diffusion model is the first stylized image obtained by the previous image generation operation.

7. The method according to claim 1, wherein: The using the first paired data to train an initial first generative adversarial network to obtain a target generative adversarial network includes: Using the first paired data to train an initial first generative adversarial network to obtain a trained first generative adversarial network; Based on a preset second original image, obtaining second pairing data through the trained first generative adversarial network; The initial second generative adversarial network is trained using the second paired data to obtain a target generative adversarial network based on the trained second generative adversarial network; wherein the network parameters of the second generative adversarial network are less than the network parameters of the first generative adversarial network.

8. The method according to claim 7, wherein: The step of obtaining second pairing data based on the preset second original image through the trained first generative adversarial network includes: Based on a preset second original image, perform multiple image generation operations through the trained first generative adversarial network to obtain a second stylized image corresponding to each image generation operation, and construct second pairing data based on the second original image and the second stylized image; Among them, the input image corresponding to the first image generation operation performed by the trained first generative adversarial network is the second original image, and the input image corresponding to the non-first image generation operation performed by the trained first generative adversarial network is the second stylized image obtained by the previous image generation operation.

9. The method according to claim 7, wherein: The performing multiple image generation operations by the trained first generative adversarial network includes: For each image generation operation, obtaining an input image corresponding to the image generation operation performed by the trained first generative adversarial network; Determine a target area in the input image corresponding to the image generation operation; The target area is stylized by using the trained first generative adversarial network to obtain a second stylized image corresponding to the image generation operation.

10. The method according to claim 9, wherein: Different image generation operations correspond to different target areas.

11. A method for constructing a stylized model, comprising: Acquire a training data set; wherein the training data set includes a target original image and a target stylized image corresponding to the target original image, and the target stylized image is obtained by using the image generation method according to any one of claims 1 to 10; The preset initial network model is trained using the training data set to construct a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model.

12. An image generating device, comprising: A model acquisition module, used to acquire a target diffusion model; A first generating module, configured to generate, based on the acquired first original image, a first stylized image corresponding to the first original image by using the target diffusion model, so as to construct first pairing data based on the first original image and the first stylized image; A network training module, configured to train an initial first generative adversarial network using the first paired data to obtain a target generative adversarial network; The second generation module is used to generate a target stylized image corresponding to the target original image by using the target generation adversarial network based on the acquired target original image.

13. A stylized model construction device, comprising: A data set acquisition module, used to acquire a training data set; wherein the training data set includes a target original image and a target stylized image corresponding to the target original image, and the target stylized image is obtained by using the image generation method according to any one of claims 1 to 10; A model building module is used to train a preset initial network model using the training data set to build a stylized model based on the trained initial network model; wherein the stylized model is used to perform style processing on its input image to obtain a stylized image corresponding to the input image of the stylized model.

14. An electronic device, comprising: a storage device having a computer program stored thereon; A processing device, used to execute the computer program in the storage device to implement the steps of the image generation method described in any one of claims 1 to 10 or the steps of the stylized model construction method described in claim 11.

15. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the steps of the image generation method according to any one of claims 1 to 10 or the steps of the stylized model construction method according to claim 11.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable medium

    CN111402112A

  • Stylized image generation method and device, computer equipment and storage medium

    CN116012488A

  • Method and system for generating content, computing equipment and storage medium

    CN117113186A

  • Personalized Machine Learning System to Edit Images Based on a Provided Style

    US20220222872A1