Image generation method based on generative adversarial network
By introducing creative fusion network and creative excitation mechanism into the image generation method and combining user feedback loops, the problem of lack of creativity and personalization of image generation in the prior art is solved, and high-quality and diverse image generation is achieved.
Patent Information
- Application Number
- CN202510551002.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing image generation methods based on generative adversarial networks cannot fully understand and blend different styles, themes and elements in user instructions, resulting in the generated images being creative and personalized, and the key features of multimodal input data cannot be effectively extracted and utilized.
An image generation method based on a generative adversarial network is adopted to collect diverse image data sets, analyze user instructions using NLP technology, build creative fusion networks that integrate different dimensions, and introduce creative excitation mechanisms and user feedback loops to optimize the quality and creativity of generated images.
Generating high-quality and creative images can better integrate user needs, improve the diversity and personalization of images, and meet users' diverse image generation needs.
Smart Images

Figure CN120070647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation method, and in particular to an image generation method based on a generative adversarial network. Background Art
[0002] Currently, most models for generating multi-style images rely on real images for style image generation. However, a few models that can generate style images from semantic maps have limitations. They can only use the image styles in the same dataset as input and cannot achieve fast migration across datasets or arbitrary styles.
[0003] The generative adversarial network (GAN) was proposed by Goodfellow et al. in 2014. It is a generative model based on deep learning and consists of a generator and a discriminator. The goal of the generator is to generate as realistic images as possible, while the discriminator attempts to distinguish between real images and generated images. The two are continuously optimized through adversarial training until the generator can generate high-quality images. Compared with other generative models, GAN can avoid complex statistical calculations and can generate higher-quality images. It performs well in the field of image generation, especially in generating realistic images, such as generating handwritten digits, face images, and artworks.
[0004] However, the existing image generation methods based on generative adversarial networks still have some deficiencies. For example, the existing methods cannot fully understand and integrate different styles, themes, and elements in user instructions, resulting in the generated images lacking creativity and personalization. In addition, the existing methods also have limitations in processing multi-modal input data and cannot effectively extract and utilize the key features in these data, thus leading to the generated images not meeting the user's requirements.
[0005] Therefore, there is an urgent need for an image generation method based on a generative adversarial network to solve the technical problems existing in the above-mentioned prior art. Summary of the Invention
[0006] The present invention overcomes the deficiencies of the prior art and provides an image generation method based on a generative adversarial network.
[0007] To achieve the above object, the technical solution adopted by the present invention is: an image generation method based on a generative adversarial network, comprising the following steps: S1. Collect an image dataset containing different styles, themes, and elements, and perform preprocessing and annotation; S2. Use NLP technology to train an instruction parsing module to parse user instructions; S3. Construct and train a creative fusion network, and according to user instructions, fuse the features of different dimensions of the image dataset to generate a creative feature representation; S4. Based on the generative adversarial network, construct a generator and a discriminator, and use adversarial training to optimize the quality and creativity of the generated images; S5. Introduce a creativity stimulation mechanism, inject random noise during the generation process, perform style transfer and variation, and establish a user feedback loop; test the generated images and optimize the generation method according to user feedback and evaluation metrics.
[0008] In a preferred embodiment of the present invention, in step S2, the parsing of the user instruction includes the following steps: S201. Receive the instruction given by the user in natural language form and parse the instruction using NLP technology; S202. Extract the key information about at least one of style, theme, and elements in the instruction; S203. Convert the extracted key information into a feature vector or label that can be understood by the machine and use it as the input of the creative fusion network.
[0009] In a preferred embodiment of the present invention, in step S3, the implementation steps of the creative fusion network include: S301. Construct a creative fusion network based on a multi-layer neural network for fusing features of different styles, themes, or elements; S302. The creative fusion network includes multiple sub-networks, which are respectively responsible for processing different types of inputs, including image style features, object shape features, and color matching features; S303. Through the attention mechanism or the feature fusion layer, organically combine features of different dimensions to form a new creative feature representation.
[0010] In a preferred embodiment of the present invention, in step S4, the generator is used to receive the feature representation output by the creative fusion network and gradually generate high-resolution images through upsampling and convolution operations; the discriminator is responsible for judging whether the generated image is real and whether it meets the requirements of the user instruction.
[0011] In a preferred embodiment of the present invention, in step S5, the implementation steps of the creativity stimulation mechanism include: S501. During the generation process of the generator, inject random noise into the input or intermediate layer of the generator to increase the diversity and unpredictability of the generated images; S502. In the creative fusion network, introduce style transfer technology to make the generated images incorporate new style elements while maintaining the original style; at the same time, perform fine-tuning on the style through variation operations; S503. Allow users to provide feedback on the generated images and adjust the generation strategy according to user feedback to optimize the generated images.
[0012] In a preferred embodiment of the present invention, a minimum loss function is provided in the generator, and the result value of the minimum loss function is optimized by means of adversarial training to obtain a high-resolution image; wherein, the formula expression of the minimum loss function is: ; In the formula, represents the expected value of the random noise sampled from the noise distribution ; represents the random noise sampled from the noise distribution ; represents the image generated by the generator; represents the discrimination result of the discriminator on the generated image; A maximum loss function is provided in the discriminator for distinguishing the image generated by the generator from the real image; wherein, the formula expression of the maximum loss function is: ; In the formula, represents the expected value of the real image sampled from the real data distribution ; represents the real image data sampled from the real data distribution ; represents that the discriminator hopes to maximize the logarithm of the discrimination result of the real image ; represents the logarithm of the minimized discrimination result of the image generated by the generator.
[0013] In a preferred embodiment of the present invention, in step S501, the formula expression of the image generated by the generator after being processed by injecting random noise is: ; wherein, represents the weight coefficient of the random noise; represents the random noise; In step S502, the style transfer technology is adopted to integrate the new style features into the generated image processed in step S501; wherein, the formula expression of the integrated generated image is: ; In the formula, represents the style transfer function; represents the new style features; represents the mutation vector.
[0014] In a preferred embodiment of the present invention, the user scores and provides feedback on the fused generated image, and incorporates the feedback value into the minimum loss function of the generator; wherein, the formula expression of the minimum loss function incorporating the feedback value is: ; in the formula, represents a hyperparameter used to control the influence degree of the user feedback value on the minimum loss function; represents the feedback value, and its range is [0, 1].
[0015] In a preferred embodiment of the present invention, an image generation system based on a generative adversarial network is proposed. Based on the above-mentioned image generation method, it includes: an instruction parsing module, a creative fusion network, and a generative adversarial network framework; wherein, the instruction parsing module is used to process and implement steps S1 - S2 to complete the parsing of user instructions; the creative fusion network receives multi-modal input data, preprocesses the input data, and converts it into feature vectors or feature maps. A feature extraction layer is set in the creative fusion network to extract key features from the input data, and convolution operations are performed on different key features using convolution kernels to generate a fused feature map; the generative adversarial network framework uses the gradient descent algorithm to iteratively optimize the parameters of the creative fusion network, and in each iteration, updates the parameters of the creative fusion network according to the gradient of the minimum loss function to obtain the minimum loss function value.
[0016] In a preferred embodiment of the present invention, before the creative fusion network receives multi-modal input data, enhancement processing such as rotation, scaling, and cropping is performed on the input data; the data types of the multi-modal input data include images, text descriptions, and audio signals; the key features include the edges, textures, and color distributions of images, the semantic information of text, and the spectral features of audio.
[0017] The present invention solves the defects existing in the background technology, and the present invention has the following beneficial effects: (1) Through the combination of the generative adversarial network and the creative fusion network, this technical solution can generate high-quality and creative images. The generator continuously optimizes the generation strategy during the training process to improve the realism and detail expressiveness of the images. At the same time, the creative fusion network can fuse various features to generate images with novelty and uniqueness, meeting the diverse needs of users.
[0018] (2) Users can interact with the system through natural language to specify features such as the style, theme, and elements of the image. The system generates corresponding images according to user instructions and continuously optimizes the generation strategy through the user feedback loop, thereby enhancing the user's sense of participation and satisfaction, and making the generated images more in line with the user's expectations and needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 is the flowchart of the method of the preferred embodiment of the present invention; Figure 2 is the flowchart of the adversarial training for optimizing the generated image of the preferred embodiment of the present invention. Detailed implementation manners
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0021] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.
[0022] As Figure 1 shown, an image generation method based on a generative adversarial network includes the following steps: S1. Collect an image dataset containing different styles, themes, and elements, and perform preprocessing and annotation; among them, the image dataset should cover a wide range of image styles (such as abstract, realistic, cartoon, etc.), themes (such as nature, city, people, etc.), and elements (such as colors, shapes, textures, etc.) to ensure the diversity and creativity of the generated images.
[0023] The preprocessing steps include image scaling, cropping, normalization, etc. to unify the image format and size for subsequent processing. And annotate the images, such as style labels, theme labels, element labels, etc., to provide a basis for subsequent instruction parsing and creative fusion.
[0024] S2. Use NLP technology to train an instruction parsing module to parse user instructions; among them, the instructions can be simple phrases or complex sentences, and the system should be able to handle instructions of different complexities.
[0025] Further, in step S2, the parsing of user instructions includes the following steps: S201. Receive the instructions given by the user in natural language form, and parse the instructions using NLP technology; S202. Extract the key information about at least one of style, theme, and elements in the instructions; when extracting key information, NLP technologies such as named entity recognition (NER) and dependency syntactic analysis can be used; Use NLP technology to train the instruction parsing module, which can parse the instructions given by the user in natural language form, extract key information and convert it into a feature vector or label that can be understood by the machine, so that the user can customize the style, theme, and elements of the generated image according to their own needs and preferences, and achieve personalized image generation.
[0026] S203. Convert the extracted key information into a feature vector or label that can be understood by the machine, and use it as the input of the creative fusion network. The conversion process is to convert the text into a vector, or directly map the text label to a predefined numerical label.
[0027] S3. Construct and train a creative fusion network, and according to the user's instructions, fuse the features of different dimensions of the image dataset to generate a creative feature representation; Further, in step S3, the implementation steps of the creative fusion network include: S301. Construct a creative fusion network based on a multi-layer neural network for fusing the features of different styles, themes, or elements; the network architecture includes convolutional layers, fully connected layers, activation functions, etc., for extracting and fusing features; S302. The creative fusion network contains multiple sub-networks, which are respectively responsible for processing different types of inputs, including image style features, object shape features, and color matching features; the sub-networks can focus on different feature dimensions respectively, such as style sub-network, theme sub-network, element sub-network, etc.; S303. Through the attention mechanism or the feature fusion layer, organically combine the features of different dimensions to form a new creative feature representation. The attention mechanism can help the network pay more attention to important features, and the feature fusion layer is responsible for combining the features of different dimensions into a new creative feature representation.
[0028] Specifically, the attention mechanism calculates the correlation between different features, assigns a weight to each feature, so that the network can pay more attention to important feature information. In the creative fusion network, a self-attention module is set up, and by calculating the similarity matrix between features, the weight of each feature is obtained, and then the weighted features are fused. Thus, the network's ability to capture key features can be enhanced, and the effect of feature fusion can be improved.
[0029] Such as Figure 2As shown in the figure, S4. Based on the generative adversarial network, a generator and a discriminator are constructed, and the quality and creativity of the generated images are optimized using adversarial training; the generator gradually converts the low-resolution feature map into a high-resolution image through upsampling (such as transposed convolution) and convolution operations.
[0030] Further, in step S4, the generator is used to receive the feature representation output by the creative fusion network and gradually generate a high-resolution image through upsampling and convolution operations; the discriminator is responsible for determining whether the generated image is real and meets the requirements of the user's instructions.
[0031] S5. Introduce a creative inspiration mechanism, inject random noise, perform style transfer and mutation during the generation process, and establish a user feedback loop; test the generated images and optimize the generation method according to user feedback and evaluation metrics.
[0032] Further, in step S5, the implementation steps of the creative inspiration mechanism include: S501. During the generation process of the generator, inject random noise into the input or intermediate layer of the generator to increase the diversity and unpredictability of the generated images; by injecting random noise, performing style transfer and mutation and other operations, increase the diversity and unpredictability of the generated images, and further meet the personalized needs of users.
[0033] Specifically, in step S501, the formula expression of the image generated by the generator after injecting random noise is: ; where represents the weight coefficient of random noise; represents random noise; S502. In the creative fusion network, introduce style transfer technology to make the generated images incorporate new style elements while maintaining the original style; at the same time, fine-tune the style through mutation operations. Specifically, in step S502, adopt style transfer technology to incorporate new style features into the generated image processed in step S501; among them, the formula expression of the fused generated image is: ; in the formula, represents the style transfer function; represents the new style feature; represents the mutation vector.
[0034] S503. Allow users to provide feedback on the generated images and adjust the generation strategy according to user feedback to optimize the generated images.
[0035] Furthermore, a minimum loss function is set in the generator, and the result value of the minimum loss function is optimized through adversarial training to obtain a high-resolution image; wherein, the formula expression of the minimum loss function is: ; In the formula, represents the expected value of the random noise sampled from the noise distribution ; represents the random noise sampled from the noise distribution ; represents the image generated by the generator; represents the discrimination result of the discriminator on the generated image.
[0036] Specifically, the following is a further description of how to continuously optimize the quality and creativity of the generated image through the alternating training of the generator and the discriminator in the generative adversarial network (GAN).
[0037] (I) Initialize network parameters: The parameters of the generator and the discriminator are initialized before the start of training, usually using the random initialization method.
[0038] (II) Alternately train the generator and the discriminator: (1) Train the discriminator: Sample a batch of real images from the real data distribution .
[0039] Sample a batch of random noise from the noise distribution , and generate the corresponding generated image through the generator.
[0040] Use the discriminator to discriminate this batch of real images and generated images, and calculate the maximum loss function of the discriminator. Use the gradient descent algorithm (such as the Adam optimizer) to update the parameters of the discriminator to minimize .
[0041] (2) Train the generator: Sample a new batch of random noise from the noise distribution . Generate the corresponding generated image through the generator. Use the discriminator to discriminate this batch of generated images, and calculate the minimum loss function of the generator. Then use the gradient descent algorithm to update the parameters of the generator to minimize .
[0042] (III) Introduce a creativity stimulation mechanism and user feedback: During each generator training process, random noise is injected into the input or intermediate layers of the generator to increase the diversity and unpredictability of the generated images.
[0043] In the creative fusion network, style transfer technology is introduced so that the generated images incorporate new style elements while maintaining the original style, and the style is fine-tuned through mutation operations.
[0044] Allow users to provide rating feedback on the generated images and incorporate the feedback value into the minimum loss function of the generator. Specifically, the feedback value (ranging from [0,1]) is multiplied by the hyperparameter and then added to the minimum loss function of the generator to adjust the generation strategy and optimize the generated images.
[0045] Here, it should be further noted that the purpose of introducing the hyperparameter is that user feedback provides a subjective evaluation of the quality of the generated images in the form of a rating . However, this subjective evaluation needs to be combined with the original loss function of the generator to form a comprehensive optimization objective. The hyperparameter plays a role in balancing these two effects. By adjusting the value of , the weight of user feedback in the optimization process can be controlled. For example, when is large, the influence of user feedback on generator training is greater, and the system will be more inclined to generate images that meet user preferences; conversely, when is small, the generator relies more on the original loss function for optimization.
[0046] Different users may have different expectations and requirements for the quality of the generated images. By adjusting the value of , the system can be adapted to the needs and preferences of different users. For example, for users who value image realism more, the value of can be appropriately reduced so that the generator relies more on the original loss function for optimization; while for users who value image creativity or personalization more, the value of can be appropriately increased so that the system pays more attention to user feedback.
[0047] During the training process, the introduction of user feedback may increase the instability of the optimization process. By introducing the hyperparameter , the influence of user feedback in the optimization process can be gradually adjusted, thereby controlling the stability and convergence of the optimization process. For example, at the beginning of training, the value of can be appropriately reduced so that the generator first performs preliminary optimization based on the original loss function; as training progresses, the value of is gradually increased.values to gradually adapt the system to the user's feedback.
[0048] (4) Adaptive adversarial loss: To balance the quality and diversity of the generated images, an adaptive adversarial loss is introduced. Metrics such as Inception Score, Fréchet Inception Distance (FID), etc. are used to evaluate the quality of the generated images. The diversity of the generated images is evaluated by calculating statistical quantities such as the entropy and mutual information of the generated image set.
[0049] According to the evaluation results of the quality and diversity of the generated images, the maximum loss function of the discriminator is dynamically adjusted. Specifically, a weighted loss function is set, where the weight coefficient changes dynamically according to the quality and diversity of the generated images.
[0050] In each round of iterative training, first calculate the quality and diversity metrics of the generated images, dynamically adjust the maximum loss function of the discriminator according to these metrics, and then use the adjusted loss function for alternating training of the discriminator and the generator.
[0051] (5) Iterative termination condition: The training process continues until a predetermined number of training rounds is reached. Alternatively, when the quality and diversity metrics of the generated images reach a certain threshold, the training can be stopped. And during the training process, the model parameters of the generator and the discriminator are regularly saved for model evaluation or deployment after the training is completed.
[0052] Preferably, in the traditional GAN training process, the loss function of the discriminator is usually fixed. However, during the training process, the quality and diversity of the generated images will continuously change. To balance the quality and diversity of the generated images, this method introduces an adaptive adversarial loss and dynamically adjusts the loss function of the discriminator according to the quality and diversity of the generated images. For example, when the quality of the generated images is low, the discriminator's error penalty for the generated images can be increased; when the diversity of the generated images is low, the discriminator's penalty for the diversity of the generated images can be increased, so that the generator can better balance the quality and diversity of the generated images during the training process.
[0053] Adaptive adversarial loss is a method of dynamically adjusting the discriminator's loss function according to the quality and diversity of the generated images. Its core idea is that during the GAN training process, the quality and diversity of the generated images are two interrelated but conflicting goals. Traditional GAN training often has difficulty optimizing these two goals simultaneously, resulting in an unstable training process and even possible problems such as mode collapse. Adaptive adversarial loss makes the generator better balance the quality and diversity of the generated images during the training process by dynamically adjusting the discriminator's loss function, thereby improving the stability of the training.
[0054] It should be further noted that the implementation steps of the adaptive adversarial loss introduced in this method are as follows: Use metrics such as Inception Score, Fréchet Inception Distance (FID), etc. to evaluate the quality of the generated images. These metrics can reflect the distance between the generated images and the real images in the feature space, thereby measuring the realism and clarity of the generated images.
[0055] Evaluate the diversity of the generated images by calculating statistics such as the entropy and mutual information of the generated image set. These metrics can reflect the degree of difference between different images in the generated image set, thereby measuring the diversity of the generated images.
[0056] According to the evaluation results of the quality and diversity of the generated images, dynamically adjust the loss function of the discriminator. Specifically, set a weighted loss function, where the weight coefficient changes dynamically according to the quality and diversity of the generated images.
[0057] When the quality of the generated images is low, increase the weight of the discriminator's error penalty for the generated images to prompt the generator to improve the quality of the generated images; when the diversity of the generated images is low, increase the weight of the discriminator's penalty for the diversity of the generated images to prompt the generator to increase the diversity of the generated images.
[0058] In each round of iterative training, first calculate the quality and diversity metrics of the generated images, dynamically adjust the loss function of the discriminator according to these metrics, and use the adjusted loss function for alternating training of the discriminator and the generator. Repeat the above process until the predetermined number of training rounds is reached or other stopping conditions are met.
[0059] The adaptive adversarial loss dynamically adjusts the loss function of the discriminator, enabling the generator to better balance the quality and diversity of the generated images during the training process, thereby helping to avoid problems such as mode collapse during the training process and improving the stability of the training.
[0060] By increasing the weight of the discriminator's penalty for the errors of the generated images, it prompts the generator to generate higher-quality images, thereby helping to improve the realism and clarity of the generated images.
[0061] By increasing the weight of the discriminator's penalty for the diversity of the generated images, it prompts the generator to generate more diverse images, thereby helping to avoid the problem of the generated images being too single or repetitive.
[0062] Specifically, the user gives a score feedback on the fused generated images and incorporates the feedback value into the minimum loss function of the generator; among them, the formula expression of the minimum loss function incorporating the feedback value is: ; In the formula, Denotes a hyperparameter used to control the influence degree of the user feedback value on the minimum loss function; Denotes a feedback value, whose range is [0, 1].
[0063] A maximum loss function is set in the discriminator to distinguish between the images generated by the generator and the real images; among them, the formula expression of the maximum loss function is: ; In the formula, Denotes the expected value of the real image sampled from the real data distribution ; Denotes the real image data sampled from the real data distribution ; Denotes the discriminator's desire to maximize the discriminant result of the real image in logarithm; Denotes the minimized discriminant result of the image generated by the generator in logarithm.
[0064] Based on the above, a further explanation is made for -E. -E is used to represent the negative value of the expectation of a random variable in the loss function of GAN. By minimizing or maximizing these expectations, the generator and the discriminator respectively optimize their own performances, so as to achieve the effect of adversarial training.
[0065] Allows users to give feedback on the generated images and adjust the generation strategy according to the user feedback. This feedback mechanism can help the system promptly discover and correct errors, and optimize the quality of the generated images.
[0066] The generative adversarial network (GAN) continuously optimizes the ability of the generator through the mutual game between the generator and the discriminator, enabling it to generate increasingly realistic images, so that the generator can capture the complex distribution characteristics in the image data and generate high-quality images.
[0067] By constructing a creative fusion network based on a multi-layer neural network and fusing the features of different styles, themes or elements, images with unique creativity can be generated, thus not only retaining the basic elements of the images, but also integrating the user's instructions and preferences, increasing the diversity of the images.
[0068] To protect user privacy, it is necessary to encrypt the user feedback data. Technologies such as symmetric encryption or asymmetric encryption are used to encrypt and store the user feedback data during transmission. At the same time, the method of anonymous feedback can be adopted to allow users to provide feedback without exposing their personal information. For example, a unique anonymous ID can be provided for each user, and this ID is associated with the user feedback data, so that useful user feedback information can be collected while protecting user privacy.
[0069] Based on the above image generation method, an image generation system based on a generative adversarial network is proposed, including: an instruction parsing module, a creative fusion network, and a generative adversarial network framework; Instruction parsing module: Instruction parsing module: Receives instructions given by the user in natural language, such as "Generate a future city scene combining classical architecture and modern technology elements".
[0070] Uses natural language processing technology (NLP) to parse the instructions and extract key information, such as style (classical and modern fusion), theme (future city), elements (architecture, technology), etc.
[0071] Converts the parsed instructions into feature vectors or labels that can be understood by machines and uses them as inputs for subsequent networks.
[0072] Creative fusion network: A multi-layer neural network used to fuse features of different styles, themes, or elements. This network can contain multiple sub-networks, each responsible for processing different types of inputs (such as image style features, object shape features, color matching features, etc.).
[0073] Through an attention mechanism or a feature fusion layer, these features in different dimensions are organically combined to form a new and creative feature representation.
[0074] Before the creative fusion network receives multi-modal input data, it performs enhancement processing such as rotation, scaling, and cropping on the input data; the data types of the multi-modal input data include images, text descriptions, and audio signals; the key features include the edges, textures, and color distributions of images, the semantic information of text, and the spectral features of audio.
[0075] Generative adversarial network (GAN) framework: Adopts a classic GAN structure, including a generator and a discriminator. The generator receives the feature representation output by the creative fusion network and gradually generates high-resolution images through operations such as upsampling and convolution. The discriminator is responsible for judging whether the generated image is real and whether it meets the requirements of the user's instructions. Through the adversarial training of the generator and the discriminator, the quality and creativity of the generated images are continuously optimized.
[0076] During the generation process, random noise is injected into the input or intermediate layers of the generator to increase the diversity and unpredictability of the generated images. In the creative fusion network, style transfer technology is introduced so that the generated images can incorporate new style elements while maintaining the original style. At the same time, through mutation operations, the style is fine-tuned to create unique visual effects.
[0077] User feedback loop: Allows users to provide feedback on the generated images, such as selecting favorite styles, elements, or suggesting improvements. The system adjusts the generation strategy based on user feedback to further optimize the generated images.
[0078] This technical solution can not only be used to generate high-quality and diverse images, but also achieve image style transfer, transferring one art style to another image, providing new possibilities for digital art creation.
[0079] By generating realistic images, it can be used for data augmentation to improve the generalization ability and performance of machine learning models. It can also be applied to fields such as virtual reality, game development, and medical image analysis, providing high-quality image generation and style transfer solutions for these fields.
[0080] Based on the inspiration of the ideal embodiments of the present invention, through the above description, relevant personnel can make various changes and modifications completely within the scope without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and must be determined according to the scope of the claims.
Claims
1. A method for generating an image based on a generative adversarial network, characterized in that: The following steps are involved: S1. Collect image datasets containing different styles, themes and elements, and perform preprocessing and annotation; S2. Use NLP technology to train the instruction parsing module to parse user instructions; S3, build and train a creative fusion network to fuse features of different dimensions of the image dataset according to user instructions and generate creative feature representations; S4. Based on the generative adversarial network, we build a generator and a discriminator, and use adversarial training to optimize the quality and creativity of the generated images. S5. Introduce a creative stimulation mechanism, inject random noise into the generation process, perform style transfer and mutation, and establish a user feedback loop; test the generated images and optimize the generation method based on user feedback and evaluation indicators.
2. The image generation method based on a generative adversarial network according to claim 1, characterized in that: In step S2, the parsing of the user instruction includes the following steps: S201, receiving instructions given by a user in natural language, and parsing the instructions using NLP technology; S202, extracting at least one key information about style, theme, and element in the instruction; S203, converting the extracted key information into a feature vector or label that can be understood by a machine, and using it as the input of the creative fusion network.
3. The image generation method based on a generative adversarial network according to claim 1, characterized in that: In step S3, the implementation steps of the creative fusion network include: S301. Construct a creative fusion network based on a multi-layer neural network to fuse features of different styles, themes or elements; S302, the creative fusion network includes multiple sub-networks, which are responsible for processing different types of inputs, including image style features, object shape features, and color matching features; S303. Through the attention mechanism or feature fusion layer, features of different dimensions are organically combined to form a new creative feature representation.
4. The image generation method based on a generative adversarial network according to claim 1, characterized in that: In step S4, the generator is used to receive the feature representation output by the creative fusion network, and gradually generate a high-resolution image through upsampling and convolution operations; the discriminator is responsible for judging whether the generated image is real and meets the requirements of the user's instructions.
5. The image generation method based on a generative adversarial network according to claim 1, characterized in that: In step S5, the implementation steps of the creativity stimulation mechanism include: S501, during the generation process of the generator, injecting random noise into the input or intermediate layer of the generator to increase the diversity and unpredictability of the generated image; S502, introducing style migration technology into the creative fusion network, so that the generated image can integrate new style elements while maintaining the original style; at the same time, fine-tuning the style through mutation operation; S503: Allow the user to provide feedback on the generated image, and adjust the generation strategy according to the user feedback to optimize the generated image.
6. The image generation method based on a generative adversarial network according to claim 4, characterized in that: The generator is provided with a minimum loss function, and the result value of the minimum loss function is optimized by adversarial training to obtain a high-resolution image; wherein the formula expression of the minimum loss function is: ; In the formula, Represents the noise distribution Random noise sampled in Expected value; Represents the noise distribution Random noise sampled in ; represents the image generated by the generator; Represents the discriminator's judgment result on the generated image; The discriminator is provided with a maximum loss function for distinguishing the image generated by the generator from the real image; wherein the formula expression of the maximum loss function is: ; In the formula, Represents the distribution of real data The real image sampled from Expected value; Represents the distribution from real data Real image data sampled from ; It means that the discriminator hopes to maximize the real image The judgment result of The logarithm of Represents the image generated by the generator The minimization result of The logarithm of .
7. The image generation method based on a generative adversarial network according to claim 5, characterized in that: In step S501, the formula expression of the image generated by the generator after the random noise injection processing is: ;in, Represents the weight coefficient of random noise; represents random noise; In step S502, the style transfer technology is used to integrate the new style features into the generated image processed in step S501; wherein the formula expression of the generated image after integration is: ; In the formula, represents the style transfer function; Indicates new stylistic features; Represents the mutation vector.
8. The image generation method based on a generative adversarial network according to claim 7, characterized in that: The user gives feedback on the generated image after fusion, and incorporates the feedback value into the minimum loss function of the generator; wherein the formula expression of the minimum loss function incorporating the feedback value is: ; In the formula, represents a hyperparameter, which is used to control the influence of the user feedback value on the minimum loss function; Represents the feedback value, which ranges from [0,1].
9. An image generation system based on a generative adversarial network, based on an image generation method according to any one of claims 1 to 8, characterized in that: include: An instruction parsing module, a creative fusion network and a generative adversarial network framework; wherein the instruction parsing module is used to process and implement steps S1-S2 to complete the parsing of user instructions; the creative fusion network receives multimodal input data, preprocesses the input data, and converts it into a feature vector or a feature map, and a feature extraction layer is set in the creative fusion network to extract key features from the input data, and use convolution kernels to perform convolution operations on different key features to generate a fused feature map; the generative adversarial network framework uses a gradient descent algorithm to iteratively optimize the parameters of the creative fusion network, and in each iteration, updates the parameters of the creative fusion network according to the gradient of the minimum loss function to obtain the minimized loss function value.
10. The image generation system based on generative adversarial network according to claim 9, characterized in that: The creative fusion network performs enhanced processing such as rotation, scaling and cropping on the multimodal input data before receiving the multimodal input data; the data types of the multimodal input data include images, text descriptions, and audio signals; the key features include edges, textures, and color distribution of images, semantic information of texts, and spectral features of audio.
Citation Information
Patent Citations
Asteroid image synthesis method and system based on generative adversarial network, and computer readable storage medium
CN113870166A
High-quality image generation method based on improved generative adversarial network
CN117095069A
Training method of generative adversarial network, and bidirectional image style conversion method and device
CN117635418A
Cloth flaw image generation method based on improved CycleGAN
CN119006367A
Deep learning-based traction type elevator steel wire rope defect detection method
CN119168948A