Image generation model training method, device, electronic device and storage medium
By acquiring and using sample roles and attribute labels to train the image generation model, the problem of single feature performance of the image generation model is solved, and a higher quality image generation effect is achieved.
Patent Information
- Application Number
- CN202111198286.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-10-14
AI Technical Summary
The feature expression dimensions of the existing image generation model are relatively single, resulting in low image generation quality and poor generation effect.
By obtaining multiple sample character images and their corresponding sample character labels and sample attribute labels, combining multiple sample attribute images, the initial image generation model is trained to obtain the target image generation model, and the image generation model's ability to express roles and attributes is enhanced.
It effectively improves the expression modeling ability of the image generation model to character and attributes, improves the image generation effect, and enhances the feature distribution and quality of the image.
Smart Images

Figure CN113947189B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, specifically to the field of artificial intelligence technologies such as deep learning and computer vision, and especially to training methods, devices, electronic devices, and storage media for image generation models. Background Art
[0002] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0003] The image generation model in the related art generates images with a relatively single feature representation dimension, resulting in low image generation quality and poor image generation effect. Summary of the Invention
[0004] The present disclosure provides a training method for an image generation model, an image generation method, an apparatus, an electronic device, a storage medium, and a computer program product.
[0005] According to a first aspect of the present disclosure, a method for training an image generation model is provided, comprising: obtaining a plurality of sample role images, the sample role images having corresponding sample role labels; determining sample attribute labels; obtaining a plurality of sample attribute images corresponding to the sample attribute labels; and training an initial image generation model based on the plurality of sample role images, the plurality of sample attribute images, the sample role labels, and the sample attribute labels to obtain a target image generation model.
[0006] According to the second aspect of the present disclosure, an image generation method is provided, comprising: obtaining character features to be processed and attribute features to be processed; inputting the character features to be processed and the attribute features to be processed into a target image generation model trained by the image generation model training method of the first aspect of the present disclosure, so as to obtain a target image output by the target image generation model.
[0007] According to a third aspect of the present disclosure, a training device for an image generation model is provided, comprising: a first acquisition module for acquiring a plurality of sample role images, the sample role images having corresponding sample role labels; a determination module for determining sample attribute labels; a second acquisition module for acquiring a plurality of sample attribute images corresponding to the sample attribute labels; and a training module for training an initial image generation model based on the plurality of sample role images, the plurality of sample attribute images, the sample role labels, and the sample attribute labels to obtain a target image generation model.
[0008] According to the fourth aspect of the present disclosure, an image generation device is provided, comprising: a third acquisition module for acquiring character features to be processed and attribute features to be processed; a generation module for inputting the character features to be processed and the attribute features to be processed into a target image generation model trained by a training device of an image generation model according to the third aspect of the present disclosure, so as to obtain a target image output by the target image generation model.
[0009] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the image generation model as described in the first aspect of the present disclosure, or to execute the image generation method as described in the second aspect of the present disclosure.
[0010] According to the sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to enable a computer to execute the training method of the image generation model as described in the first aspect of the present disclosure, or to execute the image generation method as described in the second aspect of the present disclosure.
[0011] According to the seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the training method of the image generation model as described in the first aspect of the present disclosure, or executes the steps of the image generation method as described in the second aspect of the present disclosure.
[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0014] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0015] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure;
[0020] Figure 7 is a schematic diagram according to a seventh embodiment of the present disclosure;
[0021] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the image generation model training method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure.
[0024] It should be noted that the executor of the training method of the image generation model of this embodiment is the training device of the image generation model, which can be implemented by software and / or hardware. The device can be configured in an electronic device, and the electronic device may include but is not limited to a terminal, a server, etc.
[0025] The embodiments of the present disclosure relate to the field of computer technology, and more specifically to the field of artificial intelligence technologies such as deep learning and computer vision.
[0026] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.
[0027] Deep learning involves learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.
[0028] Computer vision refers to the use of cameras and computers to replace the human eye to identify, track, and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection.
[0029] like Figure 1 As shown, the training method of the image generation model includes:
[0030] S101: Acquire multiple sample character images, where the sample character images have corresponding sample character labels.
[0031] Among them, the character images used to train the image generation model can be called sample character images, and the character images can be specifically, for example, images describing characters (wherein, characters are such as human characters, animal characters, etc.). Specifically, they can be, for example, images describing anime characters, images describing online game characters, or images describing mobile game characters, etc., without limitation.
[0032] In the disclosed embodiment, multiple sample character images can be obtained from a high-quality image library, that is, high-quality character design drawings from professional original painting designers on the game website can be obtained, and the character design drawings can be used as multiple sample character images, without limitation.
[0033] Among them, the feature label used to describe the sample role can be called a sample role label. The sample role label can be specifically, for example, a sample image feature label, a sample personality label, etc., and there is no limitation to this.
[0034] The sample image feature labels may be specifically, for example, black hair, brown eyes, etc., and the sample personality feature labels may be specifically, for example, impulsive, optimistic, etc., without limitation.
[0035] In the disclosed embodiment, after obtaining a plurality of sample character images, the plurality of sample character images may be input into a pre-trained convolutional neural network (CNN) to obtain a plurality of sample image feature labels output by the CNN.
[0036] In the disclosed embodiment, after obtaining a plurality of sample character images, the plurality of sample character images can be input into a pre-trained computer vision neural network (NNCV) and a personality diagnosis neural network (NNPD) to obtain a plurality of sample personality feature labels output by the NNCV and NNPD.
[0037] After obtaining the plurality of sample image feature labels and sample personality feature labels, the plurality of sample image feature labels and sample attribute feature labels can be used together as sample role labels.
[0038] It should be noted that, in the embodiments of the present disclosure, the acquisition process of the sample character images complies with relevant laws and regulations and does not violate public order and good morals.
[0039] S102: Determine sample attribute labels.
[0040] Among them, the feature labels used to describe some attributes can be called sample attribute labels. Sample attribute labels can be specifically, for example, sample scene labels, sample clothing labels, sample prop labels, etc., and there is no limitation on this.
[0041] The attribute may be related to the above-mentioned character, and may reflect the scene, clothing, and type and characteristics of the character, and there is no restriction on this.
[0042] For example, to determine the sample attribute labels, you can determine sample scene labels such as pure white snow scene, dark forest, sample clothing labels such as golden armor, black cloak, sample prop labels such as silver axe, long sword, and other sample attribute labels, and there is no restriction on this.
[0043] In the disclosed embodiment, the plurality of sample character images may have various attributes such as scenes, costumes, and props, and thus corresponding sample attribute labels may be determined based on the various attributes of the sample character images.
[0044] For example, determining the sample attribute labels may be inputting sample character images with multiple attributes into a pre-trained convolutional neural network (CNN) model to obtain multiple sample attribute labels output by the CNN model, and there is no limitation to this.
[0045] S103: Acquire multiple sample attribute images corresponding to the sample attribute labels.
[0046] Among them, the attribute images used to train the image generation model can be called sample attribute images. The sample attribute images can be used to describe the attributes of the sample, such as the scene, clothing, and props. The sample attribute images and sample character images can together constitute multiple sample images.
[0047] That is, in the embodiment of the present disclosure, the sample image can be divided into a human role area and an attribute area, and then the sample image corresponding to the human role area can be used as a sample character image, and the image corresponding to the attribute area can be used as a sample attribute image.
[0048] After determining the sample attribute label as mentioned above, multiple sample attribute images corresponding to the sample attribute label can be obtained. When the sample attribute label is a sample scene label, a sample clothing label, and a sample prop label, the sample attribute image can be specifically, for example, a sample prop image, a sample clothing image, and a sample scene image, etc., and there is no restriction on this.
[0049] In some embodiments, obtaining multiple sample attribute images corresponding to the sample attribute label can be performed by searching among multiple sample attribute images based on the sample attribute label to determine the sample attribute image corresponding to the sample attribute label, and there is no limitation on this. Alternatively, any other possible method can be used to obtain multiple sample attribute images corresponding to the sample attribute label, such as a feature matching method, a model prediction method, etc., and there is no limitation on this.
[0050] For example, when the sample attribute label is a sample scene label, obtaining a plurality of sample attribute images corresponding to the sample attribute label may be obtaining a sample scene image corresponding to the sample scene label from the plurality of sample attribute images.
[0051] Optionally, in other embodiments, obtaining multiple sample attribute images corresponding to the sample attribute label can be performed by respectively parsing the multiple sample attribute images corresponding to the sample attribute label from the multiple sample images. Since the multiple sample attribute images corresponding to the sample attribute label are parsed from the multiple sample images, the efficiency of obtaining the sample attribute images is effectively improved, and the sample images can be some original images that have been formed in a high-quality image library. Therefore, while ensuring the efficiency of obtaining the sample attribute images, the quality of the sample attribute images can be effectively guaranteed, and the sample attribute images and the sample attribute labels have better adaptability, thereby effectively ensuring the reliability of the sample attribute images.
[0052] That is, in the embodiment of the present disclosure, after obtaining the sample attribute label, multiple sample images can be parsed according to the sample attribute label to obtain multiple sample attribute images corresponding to the sample attribute label.
[0053] S104: Train an initial image generation model based on the multiple sample character images, the multiple sample attribute images, the sample character labels, and the sample attribute labels to obtain a target image generation model.
[0054] Among them, the image generation model obtained in the initial stage of training can be called the initial image generation model. The initial image generation model can be an artificial intelligence model, such as a neural network model or a machine learning model. Of course, any other possible artificial intelligence model that can perform image generation tasks can also be used, without limitation.
[0055] In the disclosed embodiment, the initial image generation model can be specifically, for example, a Conditional Generative Adversarial Networks (CGAN) model, that is, the CGAN model can be trained based on multiple sample character images, multiple sample attribute images, sample character labels, and sample attribute labels, and the trained CGAN model can be used as the target image generation model, without limitation.
[0056] After obtaining multiple sample attribute images corresponding to the sample attribute labels, the initial image generation model can be trained based on the multiple sample roles, multiple sample attribute images, sample role labels, and sample attribute labels to obtain a trained image generation model, which can be called a target image generation model.
[0057] In some embodiments, multiple sample roles, multiple sample attribute images, sample role labels, and sample attribute labels can be input into the initial image generation model to iteratively train the initial image generation model until the trained image generation model meets certain convergence conditions, and the image generation model that meets the convergence conditions is used as the target image generation model.
[0058] For example, a loss function can be pre-configured for the initial image generation model. During the training of the initial image generation model, multiple sample roles, multiple sample attribute images, sample role labels, and sample attribute labels are input into the initial image generation model to obtain the prediction processing results of the output of the image generation model. The prediction processing results and the annotation processing results (the annotation processing results can be used to assist in determining whether the model has converged) can then be used as input parameters of the loss function, and the loss value output by the loss function can be determined. The loss value is then compared with the set loss threshold to determine whether the convergence timing is met (if the convergence timing is met, it can be indicated that the model has converged). If the model is determined to have converged, the trained image generation model can be used as the target image generation model.
[0059] In this embodiment, by obtaining multiple sample role images, the sample role images have corresponding sample role labels, and the sample attribute labels are determined, and then multiple sample attribute images corresponding to the sample attribute labels are obtained, and the initial image generation model is trained based on the multiple sample role images, multiple sample attribute images, sample role labels, and sample attribute labels to obtain the target image generation model, which can effectively assist in improving the expression modeling ability of the target image generation model for sample attributes and sample roles. When the target image generation model is used to generate the target image, the target image can characterize the feature distribution of the role and attribute dimensions, effectively improving the image generation effect, and improving the feature modeling effect of the target image generated by the target image generation model.
[0060] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure.
[0061] like Figure 2 As shown, the training method of the image generation model includes:
[0062] S201: Acquire multiple sample character images, where the sample character images have corresponding sample character labels.
[0063] S202: Determine sample attribute labels.
[0064] S203: Acquire multiple sample attribute images corresponding to the sample attribute labels.
[0065] The description of S201 - S203 can be found in the above embodiment and will not be repeated here.
[0066] S204: Analyze and obtain multiple sample role features from multiple sample role images according to the sample role labels.
[0067] After determining the sample role labels corresponding to the sample role images, multiple sample role features can be obtained by parsing the multiple sample role images according to the sample role labels.
[0068] Optionally, in other embodiments, multiple sample role features are respectively parsed from multiple sample role images based on the sample role labels. Sample image features and sample personality features can be parsed from sample role images based on the sample role labels, and multiple sample image features and multiple sample personality features are collectively used as multiple sample role features. Since multiple sample image features and multiple sample personality features are collectively used as multiple sample role features, the sample role features can represent the features of the image dimension and the personality dimension, thereby effectively expanding the representation dimension of the sample role features, avoiding the simplification of the sample role features, and effectively improving the diversity of the sample role features. While effectively ensuring the diversity of the sample role features, it can effectively improve the richness and representation ability of the sample role features. When the image generation model is assisted in training based on multiple sample role features, the trained target image generation model can generate a target image that represents the image features and personality features, thereby helping to improve the generation effect of the target image.
[0069] Among them, the features used to describe the sample character image can be called sample image features, and the sample image features can be specifically, for example, sample facial features, sample face shape features, sample hair features, etc., without limitation.
[0070] Among them, the characteristics used to describe the personality of the sample character can be called sample personality characteristics, and the sample personality characteristics can be specifically, for example, happy personality characteristics, depressed personality characteristics, angry personality characteristics, etc., without limitation.
[0071] In the disclosed embodiment, after obtaining the sample role label, the sample image features and the sample personality features can be analyzed from the sample role image according to the sample role label.
[0072] For example, if the sample character labels obtained are black eyes and a small mouth, the sample facial features can be parsed from the sample image based on the sample character labels. If the sample character labels obtained are crying and depressed, the depressed personality characteristics can be parsed from the sample image based on the sample character labels. There is no restriction on this.
[0073] S205: According to the sample attribute labels, parse and obtain a plurality of sample attribute features from the plurality of sample attribute images respectively.
[0074] After the sample attribute labels are determined, multiple sample attribute features can be obtained by parsing the multiple sample attribute images according to the sample attribute labels.
[0075] Optionally, in other embodiments, according to the sample attribute labels, multiple sample attribute features are respectively parsed from multiple sample attribute images. It can be that according to the sample attribute labels, sample scene features, sample clothing features, and sample prop features are parsed from the sample attribute images, and multiple sample scene features, multiple sample clothing features, and multiple sample prop features are collectively used as multiple sample attribute features. Since the sample clothing features, sample scene features and sample prop features are collectively used as multiple sample attribute features, the sample attribute features can characterize the features of the scene dimension, clothing dimension and prop dimension, thereby effectively expanding the representation dimension of the sample attribute features, avoiding the simplification of the sample attribute features, and effectively improving the diversity of the sample attribute features. While effectively ensuring the diversity of the sample attribute features, it can effectively improve the richness and representation ability of the sample attribute features. When the image generation model is assisted in training based on multiple sample attribute features, the trained target image generation model can generate a target image that characterizes the features of the scene dimension, clothing dimension and prop dimension, thereby helping to improve the generation effect of the target image.
[0076] The features used to describe the sample scene may be referred to as sample scene features, and the sample scene features may be, for example, snowy scene features or forest scene features, without limitation.
[0077] The features used to describe the sample clothing may be referred to as sample clothing features, and the sample clothing features may specifically be, for example, clothing features of skirts or clothing features of pants, without limitation.
[0078] Among them, the features used to describe the sample props can be called sample prop features, and the sample prop features can be specifically, for example, the prop features of holding a knife or the prop features of holding a sword, without limitation.
[0079] In the embodiment of the present disclosure, after the sample attribute label is determined, multiple sample attribute features can be obtained by respectively parsing the multiple sample attribute images according to the sample attribute label.
[0080] For example, if the obtained sample attribute label is a black forest, the sample scene feature of the forest scene can be parsed from the sample image according to the sample attribute label.
[0081] In the embodiment of the present disclosure, after obtaining multiple sample attribute features, a minority of sample attribute features can be eliminated from the multiple sample attribute features (a minority of sample attribute features refers to features that have a small number of samples to which the features belong due to uneven distribution of sample attributes), and then the other attribute features after elimination can be used as multiple sample attribute features, thereby effectively ensuring the image generation quality of the image generation model.
[0082] S206: Training an initial image generation model based on the multiple sample character images, the multiple sample attribute images, the multiple sample character features, and the multiple sample attribute features to obtain a target image generation model.
[0083] After obtaining the sample role features and sample attribute features, the initial image generation model can be trained based on the sample role features, sample attribute features, multiple sample role images, and multiple sample attribute images to obtain the target image generation model. In this way, both the image of the role dimension and the image of the attribute dimension are used as sample images to assist in training the image generation model, thereby effectively enriching the training data set of the image generation model and ensuring the modeling learning effect of the trained target image generation model on the image of the role dimension and the image of the attribute dimension, thereby improving the training effect of the target image generation model, and assisting in expanding the application scenarios of the target image generation model and improving the application value of the target image generation model.
[0084] Optionally, in some embodiments, as Figure 3 As shown, Figure 3 is a schematic diagram according to a third embodiment of the present disclosure, wherein an initial image generation model is trained based on multiple sample character images, multiple sample attribute images, multiple sample character features, and multiple sample attribute features to obtain a target image generation model, including:
[0085] S301: Inputting multiple sample character features into an initial character parsing model respectively to obtain multiple predicted character images output by the character parsing model.
[0086] Among them, the initial role parsing model can be used to generate corresponding role images based on multiple sample role features.
[0087] The initial role parsing model may be an artificial intelligence model, specifically a neural network model or a machine learning model, for example, without limitation.
[0088] That is, multiple sample character features can be input into the initial character parsing model to obtain multiple character images output by the character parsing model. The character images can be called predicted character images.
[0089] S302: Inputting multiple sample attribute features into the initial preprocessing model respectively to obtain multiple predicted attribute images output by the preprocessing model.
[0090] Among them, the initial preprocessing model can be used to generate corresponding attribute images according to multiple sample attribute features.
[0091] The initial preprocessing model may be an artificial intelligence model, specifically a neural network model or a machine learning model, for example, without limitation.
[0092] That is, multiple sample attribute features can be input into the initial preprocessing model to obtain multiple attribute images output by the preprocessing model. The attribute images can be called predicted attribute images.
[0093] S303: Training an image generation model to be trained based on multiple sample character images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted character images to obtain a target image generation model.
[0094] Optionally, in some embodiments, the image generation model to be trained is trained based on multiple sample role images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted role images to obtain a target image generation model. Multiple sample synthetic images can be generated, and the sample synthetic image is synthesized by a first sample role image and a first sample attribute image. The first sample role image belongs to multiple sample role images, and the first sample attribute image belongs to multiple sample attribute images. The first predicted role feature corresponding to the first sample role image is determined, and the first predicted attribute feature corresponding to the first sample attribute image is determined. The first sample attribute image and the first predicted attribute feature are then input into the image generation model to be trained to obtain a predicted synthetic image output by the image generation model. When the prediction loss value between the predicted synthetic image and the sample synthetic image meets the set conditions, the trained image generation model is used as the target image generation model. Since the initial image generation model is trained in combination with the loss function, the convergence timing of the model can be accurately judged, thereby avoiding the influence of other factors on the accuracy of the convergence timing judgment, and can effectively improve the accuracy of the convergence timing judgment, and effectively improve the training effect of the image generation model.
[0095] Among them, any sample character image among the multiple sample character images can be called a first sample character image, and accordingly, the character feature corresponding to the first sample character image can be called a first predicted character feature.
[0096] Any sample attribute image among the multiple sample attribute images can be referred to as a first sample attribute image, and accordingly, the attribute feature corresponding to the first sample attribute image can be referred to as a first predicted attribute feature.
[0097] In the embodiment of the present disclosure, the first sample character image and the first sample attribute image can be synthesized to obtain a synthesized image, which can be called a sample synthesized image. The sample synthesized image can be used to assist in determining the convergence timing of the model during the training process of the image generation model.
[0098] In the embodiment of the present disclosure, a feature extraction method can be used to determine the first predicted character feature corresponding to the first sample character image, and a feature extraction method can be used to determine the first predicted attribute feature corresponding to the first sample attribute image, without limitation.
[0099] In an embodiment of the present disclosure, after determining the first predicted character feature corresponding to the first sample character image and determining the first predicted attribute feature corresponding to the first sample attribute image, the first sample attribute image and the first predicted attribute feature can be input into the image generation model to be trained to obtain an image output by the image generation model, which can be called a predicted synthetic image.
[0100] Among them, the predicted synthetic image and the sample synthetic image can be used together to determine whether the image generation model converges.
[0101] In an embodiment of the present disclosure, a loss function can be pre-configured for an image generation model. During the training process of the image generation model, the predicted synthetic image and the sample synthetic image can be used as input parameters of the loss function, and the predicted loss value between the predicted synthetic image and the sample synthetic image output by the loss function can be determined. If the predicted loss value between the predicted synthetic image and the sample synthetic image meets the set conditions (the set conditions can be adaptively configured according to the actual business scenario), the trained image generation model will be used as the target image generation model.
[0102] Of course, any other possible method can also be used to train the image generation model to be trained based on multiple sample character images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted character images to obtain the target image generation model, and there is no limitation on this.
[0103] For example, multiple sample character images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted character images can be input into the image generation model to be trained to obtain the prediction processing results of the output of the image generation model. The prediction processing results and the annotation processing results (the annotation processing results can be used to assist in determining whether the model has converged) can then be used as input parameters of the loss function, and the loss value output by the loss function can be determined. The loss value is then compared with the set loss threshold to determine whether the convergence timing is met (if the convergence timing is met, it can be indicated that the model has converged). If the model is determined to have converged, the trained image generation model can be used as the target image generation model.
[0104] In the disclosed embodiment, multiple sample role features are respectively input into the initial role parsing model to obtain multiple predicted role images output by the role parsing model, and multiple sample attribute features are respectively input into the initial preprocessing model to obtain multiple predicted attribute images output by the preprocessing model, and the image generation model to be trained is trained based on the multiple sample role images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted role images to obtain a target image generation model. Since the model is trained based on the multiple sample role images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted role images, the trained target image generation model can adaptively generate images of the role dimension and images of the attribute dimension, and can effectively guarantee the quality of the generated images of the role dimension and images of the attribute dimension, and can effectively reduce the manpower cost required to generate images of the role dimension and images of the attribute dimension, thereby greatly enriching the training and application scenarios of the target image generation model.
[0105] In this embodiment, by obtaining multiple sample role images, the sample role images have corresponding sample role labels, and the sample attribute labels are determined, and then multiple sample attribute images corresponding to the sample attribute labels are obtained, and according to the sample role labels, multiple sample role features are respectively parsed from the multiple sample role images, and according to the sample attribute labels, multiple sample attribute features are respectively parsed from the multiple sample attribute images. After obtaining the sample role features and sample attribute features, the initial image generation model can be trained according to the sample role features, the sample attribute features, the multiple sample role images and the multiple sample attribute images to obtain the target image generation model. In this way, both the images of the role dimension and the images of the attribute dimension are used as sample images to assist in training the image generation model, thereby effectively enriching the training data set of the image generation model, ensuring the modeling learning effect of the trained target image generation model on the images of the role dimension and the images of the attribute dimension, thereby improving the training effect of the target image generation model, and assisting in expanding the application scenarios of the target image generation model and improving the application value of the target image generation model.
[0106] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure.
[0107] It should be noted that the executor of the image generation method of this embodiment is an image generation device, which can be implemented by software and / or hardware. The device can be configured in an electronic device, which may include but is not limited to a terminal, a server, etc.
[0108] like Figure 4 As shown, the image generation method includes:
[0109] S401: Obtaining the role features and attribute features to be processed.
[0110] Among them, the character features currently to be processed for image generation can be called to-be-processed character features, and the attribute features currently to be processed for image generation can be called to-be-processed attribute features.
[0111] The character features to be processed may include, for example, facial features, face shape features, pleasant personality features, etc., without limitation.
[0112] The attribute features to be processed may be, for example, snowy scene features, skirt clothing features, knife-holding prop features, etc., without limitation.
[0113] Among them, the character features to be processed and the attribute features to be processed can be adaptively input by the original painting designer, that is, a feature input interface can be configured for the image generation device, and the character features to be processed and the attribute features to be processed entered by the original painting designer can be received through the interface. Alternatively, any other possible method can be used to obtain the character features to be processed and the attribute features to be processed, and there is no limitation on this.
[0114] S402: Inputting the character features to be processed and the attribute features to be processed into the target image generation model trained by the above-mentioned image generation model training method to obtain the target image output by the target image generation model.
[0115] After obtaining the character features to be processed and the attribute features to be processed, the character features to be processed and the attribute features to be processed can be input into the target image generation model trained by the training method of the image generation model as mentioned above to obtain the image output by the target image generation model, which can be called the target image.
[0116] In this embodiment, by obtaining the role features to be processed and the attribute features to be processed, and inputting the role features to be processed and the attribute features to be processed into the target image generation model trained by the training method of the image generation model as described above, the target image output by the image generation model is obtained. Since the target image generation model is trained based on the sample attribute labels and the sample role labels, when the trained image generation model is used to generate an image, the image can characterize the feature distribution of the role and attribute dimensions, thereby effectively improving the image generation quality and effectively improving the image generation effect of the image generation model.
[0117] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure.
[0118] like Figure 5 As shown, the training device 50 of the image generation model includes:
[0119] A first acquisition module 501 is used to acquire a plurality of sample character images, each of which has a corresponding sample character label;
[0120] Determination module 502, used to determine the sample attribute label;
[0121] The second acquisition module 503 is configured to acquire a plurality of sample attribute images corresponding to the sample attribute labels; and
[0122] The training module 504 is used to train the initial image generation model based on multiple sample character images, multiple sample attribute images, sample character labels, and sample attribute labels to obtain a target image generation model.
[0123] In some embodiments of the present disclosure, Figure 6 As shown, Figure 6 6 is a schematic diagram of a sixth embodiment of the present disclosure. The image generation model training device 60 includes: a first acquisition module 601, a determination module 602, a second acquisition module 603, and a training module 604. The second acquisition module 603 is specifically configured to:
[0124] A plurality of sample attribute images corresponding to the sample attribute labels are respectively obtained by parsing the plurality of sample images.
[0125] In some embodiments of the present disclosure, the training module 604 includes:
[0126] A first parsing submodule 6041 is configured to parse a plurality of sample character images according to the sample character labels to obtain a plurality of sample character features;
[0127] The second parsing submodule 6042 is configured to parse the plurality of sample attribute images to obtain a plurality of sample attribute features according to the sample attribute labels;
[0128] The training submodule 6043 is used to train the initial image generation model based on multiple sample character images, multiple sample attribute images, multiple sample character features, and multiple sample attribute features to obtain a target image generation model.
[0129] In some embodiments of the present disclosure, the initial image generation model includes: an initial character parsing model, an initial preprocessing model, and an image generation model to be trained;
[0130] The training submodule 6043 includes:
[0131] A first generating unit 60431 is configured to input a plurality of sample character features into an initial character parsing model to obtain a plurality of predicted character images output by the character parsing model;
[0132] The second generating unit 60432 is used to input the multiple sample attribute features into the initial preprocessing model respectively to obtain multiple predicted attribute images output by the preprocessing model;
[0133] The training unit 60433 is used to train the image generation model to be trained based on multiple sample character images, multiple sample attribute images, multiple predicted attribute images, and multiple predicted character images to obtain a target image generation model.
[0134] In some embodiments of the present disclosure, the training unit 60433 is specifically configured to:
[0135] Generate multiple sample composite images, where the sample composite images are synthesized by a first sample character image and a first sample attribute image, where the first sample character image belongs to multiple sample character images, and the first sample attribute image belongs to multiple sample attribute images;
[0136] Determining a first predicted character feature corresponding to the first sample character image, and determining a first predicted attribute feature corresponding to the first sample attribute image;
[0137] Inputting the first sample attribute image and the first predicted attribute feature into the image generation model to be trained to obtain a predicted composite image output by the image generation model;
[0138] If the prediction loss value between the predicted synthetic image and the sample synthetic image meets the set conditions, the trained image generation model is used as the target image generation model.
[0139] In some embodiments of the present disclosure, the first parsing submodule 6041 is specifically configured to:
[0140] According to the sample role label, the sample image features and sample personality features are obtained from the sample role image;
[0141] Multiple sample image features and multiple sample personality features are taken together as multiple sample role features.
[0142] In some embodiments of the present disclosure, the second parsing submodule 6042 is specifically configured to:
[0143] According to the sample attribute labels, the sample scene features, sample clothing features, and sample prop features are parsed from the sample attribute images;
[0144] Multiple sample scene features, multiple sample clothing features, and multiple sample prop features are collectively used as multiple sample attribute features.
[0145] It is understandable that the present embodiment Figure 6The training device 60 of the image generation model in the embodiment may have the same function and structure as the training device 50 of the image generation model in the above embodiment, the first acquisition module 601 may have the same function and structure as the first acquisition module 501 in the above embodiment, the determination module 602 may have the same function and structure as the determination module 502 in the above embodiment, the second acquisition module 603 may have the same function and structure as the second acquisition module 503 in the above embodiment, and the training module 604 may have the same function and structure as the training module 504 in the above embodiment.
[0146] It should be noted that the aforementioned explanation of the training method of the image generation model is also applicable to the training device of the image generation model in this embodiment.
[0147] In this embodiment, by obtaining multiple sample role images, the sample role images have corresponding sample role labels, and the sample attribute labels are determined, and then multiple sample attribute images corresponding to the sample attribute labels are obtained, and the initial image generation model is trained based on the multiple sample role images, multiple sample attribute images, sample role labels, and sample attribute labels to obtain the target image generation model, which can effectively assist in improving the expression modeling ability of the target image generation model for sample attributes and sample roles. When the target image generation model is used to generate the target image, the target image can characterize the feature distribution of the role and attribute dimensions, effectively improving the image generation effect, and improving the feature modeling effect of the target image generated by the target image generation model.
[0148] Figure 7 is a schematic diagram according to a seventh embodiment of the present disclosure.
[0149] like Figure 7 As shown, the image generating device 70 includes:
[0150] The third acquisition module 701 is used to acquire the role characteristics and attribute characteristics to be processed;
[0151] The generation module 702 is used to input the character features to be processed and the attribute features to be processed into the target image generation model trained by the training device of the above-mentioned image generation model to obtain the target image output by the target image generation model.
[0152] It should be noted that the aforementioned explanation of the image generating method is also applicable to the image generating device of this embodiment and will not be repeated here.
[0153] In this embodiment, by obtaining the role features to be processed and the attribute features to be processed, and inputting the role features to be processed and the attribute features to be processed into the target image generation model trained by the training method of the image generation model as described above, the target image output by the image generation model is obtained. Since the target image generation model is trained based on the sample attribute labels and the sample role labels, when the trained image generation model is used to generate an image, the image can characterize the feature distribution of the role and attribute dimensions, thereby effectively improving the image generation quality and effectively improving the image generation effect of the image generation model.
[0154] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0155] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the training method of the image generation model of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0156] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0157] Multiple components in device 800 are connected to I / O interface 805, including: a generation unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0158] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method of the image generation model, or the image generation method. For example, in some embodiments, the training method of the image generation model, or the image generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, the training method of the image generation model described above, or one or more steps of the image generation method, can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the image generation model training method or the image generation method in any other appropriate manner (for example, by means of firmware).
[0159] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0160] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0161] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0163] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0164] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0165] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0166] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A training method for an image generation model, comprising: Acquire a plurality of sample character images, each of the sample character images having corresponding sample character labels, wherein the sample character labels include a sample image feature label and a sample personality label; Determining sample attribute labels, wherein the sample attribute labels include sample scene labels, sample clothing labels, and sample prop labels; Acquire a plurality of sample attribute images corresponding to the sample attribute labels; as well as Training an initial image generation model based on the multiple sample character images, the multiple sample attribute images, the sample character labels, and the sample attribute labels to obtain a target image generation model, wherein the initial image generation model is a conditional generative adversarial network model; The initial image generation model includes: an initial role parsing model, an initial preprocessing model, and an image generation model to be trained. The initial image generation model is trained based on the multiple sample role images, the multiple sample attribute images, the sample role labels, and the sample attribute labels to obtain a target image generation model, including: According to the sample role labels, respectively analyzing the plurality of sample role images to obtain a plurality of sample role features; According to the sample attribute labels, respectively analyzing the plurality of sample attribute images to obtain a plurality of sample attribute features; Inputting the plurality of sample character features into the initial character parsing model respectively to obtain a plurality of predicted character images output by the character parsing model; Inputting the plurality of sample attribute features into the initial preprocessing model respectively to obtain a plurality of predicted attribute images output by the preprocessing model; Training the image generation model to be trained based on the multiple sample character images, the multiple sample attribute images, the multiple predicted attribute images, and the multiple predicted character images to obtain the target image generation model; The method further comprises: Inputting the plurality of sample character images into a pre-trained convolutional neural network model to obtain a plurality of sample image feature labels output by the convolutional neural network model; Determining the sample attribute label includes: Sample character images with multiple attributes are input into a pre-trained convolutional neural network model to obtain multiple sample attribute labels output by the convolutional neural network model.
2. The method according to claim 1, wherein the plurality of sample character images respectively belong to a plurality of corresponding sample images, wherein: The acquiring of a plurality of sample attribute images corresponding to the sample attribute labels includes: A plurality of sample attribute images corresponding to the sample attribute labels are respectively obtained by parsing the plurality of sample images.
3. The method according to claim 1, wherein The step of respectively parsing the plurality of sample character images to obtain a plurality of sample character features according to the sample character labels includes: Analyzing the sample character image to obtain sample image features and sample personality features according to the sample character label; The plurality of sample image features and the plurality of sample personality features are collectively used as the plurality of sample role features.
4. The method according to claim 1, wherein The step of respectively parsing the plurality of sample attribute images to obtain a plurality of sample attribute features according to the sample attribute labels includes: According to the sample attribute label, sample scene features, sample clothing features, and sample prop features are parsed from the sample attribute image; The plurality of sample scene features, the plurality of sample clothing features, and the plurality of sample prop features are collectively used as the plurality of sample attribute features.
5. A method for generating an image, comprising: Obtaining the role features and attribute features to be processed; The character features to be processed and the attribute features to be processed are input into a target image generation model trained by the image generation model training method as described in any one of claims 1 to 4 above, so as to obtain a target image output by the target image generation model.
6. A training device for an image generation model, comprising: A first acquisition module is configured to acquire a plurality of sample character images, wherein the sample character images have corresponding sample character labels, and the sample character labels include sample image feature labels and sample personality labels; A determination module, configured to determine a sample attribute label, wherein the sample attribute label includes a sample scene label, a sample clothing label, and a sample prop label; A second acquisition module is used to acquire a plurality of sample attribute images corresponding to the sample attribute labels; as well as a training module, configured to train an initial image generation model based on the plurality of sample character images, the plurality of sample attribute images, the sample character labels, and the sample attribute labels to obtain a target image generation model, wherein the initial image generation model is a conditional generative adversarial network model; The initial image generation model includes: an initial character parsing model, an initial preprocessing model, and an image generation model to be trained; Wherein, the training module includes: A first parsing submodule is configured to parse the plurality of sample character images according to the sample character labels to obtain a plurality of sample character features; A second parsing submodule is configured to parse the plurality of sample attribute images to obtain a plurality of sample attribute features according to the sample attribute labels; a training submodule, configured to train an initial image generation model based on the plurality of sample character images, the plurality of sample attribute images, the plurality of sample character features, and the plurality of sample attribute features to obtain a target image generation model; The training submodule includes: a first generating unit, configured to input the plurality of sample character features into the initial character parsing model respectively, to obtain a plurality of predicted character images output by the character parsing model; a second generating unit, configured to input the plurality of sample attribute features into the initial preprocessing model respectively, to obtain a plurality of predicted attribute images output by the preprocessing model; a training unit, configured to train the image generation model to be trained based on the plurality of sample character images, the plurality of sample attribute images, the plurality of predicted attribute images, and the plurality of predicted character images to obtain the target image generation model; The device also performs: Inputting the plurality of sample character images into a pre-trained convolutional neural network model to obtain a plurality of sample image feature labels output by the convolutional neural network model; The determining module is specifically configured to: Sample character images with multiple attributes are input into a pre-trained convolutional neural network model to obtain multiple sample attribute labels output by the convolutional neural network model.
7. The apparatus according to claim 6, wherein the plurality of sample character images respectively belong to a corresponding plurality of sample images, wherein: The second acquisition module is specifically configured to: A plurality of sample attribute images corresponding to the sample attribute labels are respectively obtained by parsing the plurality of sample images.
8. The device according to claim 6, wherein The first parsing submodule is specifically used to: Analyzing the sample character image to obtain sample image features and sample personality features according to the sample character label; The plurality of sample image features and the plurality of sample personality features are collectively used as the plurality of sample role features.
9. The device according to claim 6, wherein The second parsing submodule is specifically used to: According to the sample attribute label, sample scene features, sample clothing features, and sample prop features are parsed from the sample attribute image; The plurality of sample scene features, the plurality of sample clothing features, and the plurality of sample prop features are collectively used as the plurality of sample attribute features.
10. An image generating device, comprising: The third acquisition module is used to obtain the role characteristics and attribute characteristics to be processed; A generation module is used to input the character features to be processed and the attribute features to be processed into a target image generation model trained by the training device of the image generation model as described in any one of claims 6 to 9 above, so as to obtain a target image output by the target image generation model.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4, or the method according to claim 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 4, or to execute the method according to claim 5.
13. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 4, or performs the steps of the method according to claim 5.
Citation Information
Patent Citations
Virtual image generation method, device and equipment, and readable storage medium
CN108510437A
An image generation method and a terminal device
CN109215007A