Image intelligent generation method and system based on AIGC
By combining a variety of advanced technical models, an image generation model is constructed and optimized and content detection is carried out during the generation process, problems such as missing details, unnaturalness, and insufficient generalization capabilities of image generation in the existing technology are solved, and high-quality, strong sense of reality and compliant image generation effects are achieved.
Patent Information
- Application Number
- CN202510090664.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AIGC-based image intelligent generation technology faces problems such as lack of details or unnatural results in lack of realism, insufficient generalization capabilities of model, complex cross-modal generation, violations or biases in training data, and the generated content does not comply with legal norms and social norms.
By effectively combining various technologies such as Transformer model, diffusion model, generative adversarial network model and variational autoencoder, an image generation model is constructed, and image optimization and content detection are performed during the generation process to ensure the quality, interpretability and security of the generated images.
It realizes the generation of high-quality, realistic and standardized images, improves the generalization ability and cross-modal generation ability of the model, and ensures the security and compliance of the generated content.
Smart Images

Figure CN120014088A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation, and in particular relates to an AIGC-based image intelligent generation method and system. Background Art
[0002] AIGC (Artificial Intelligence-Generated Content) refers to content generated by artificial intelligence, covering various forms of creation such as text, images, audio, video, etc. AIGC uses advanced artificial intelligence technologies, especially natural language processing (NLP), computer vision (CV) and generative models (such as generative adversarial networks GANs, diffusion models, etc.) to create innovative, authentic and diverse content.
[0003] AIGC-based intelligent image generation is the process of generating image content using artificial intelligence technology. By training deep learning models, especially generative adversarial networks (GANs), diffusion models, and variational autoencoders (VAE), AI can generate realistic or creative images based on given inputs (such as text descriptions, existing images, etc.). Although AIGC-based intelligent image generation has shown great potential in many fields, it still faces many challenges, such as lack of realism due to missing or unnatural details, the generalization ability of the model needs to be improved, cross-modal generation (such as text-generated images) is complex, and a large amount of data used for training is illegal or biased and the generated content does not comply with legal regulations and social norms. Therefore, in order to achieve wider application and sustainable development, further improvements must be made to the above content. At the same time, improving the quality, interpretability and security of AI-generated images is also the key to solving current problems. Summary of the invention
[0004] The purpose of the present invention is to provide an AIGC-based image intelligent generation method and system, which can be implemented by the following technical solutions: In a first aspect, an embodiment of the present application provides an AIGC-based image intelligent generation method, comprising the following steps: Build an image generation model; Classifying the data input into the image generation model into text input data and image input data; Processing the text input data and the image input data to obtain text processed data and image processed data respectively; Inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Performing content detection on the model-optimized image to determine whether the image content complies with legal regulations and outputting the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
[0005] Preferably, the constructing of the image generation model comprises: Construct a text-to-image model based on the Tansformer model and diffusion model; Construct an image generation model based on the generative adversarial network model and the diffusion model; The text generation image model and the image generation image model are fused to generate the image generation model.
[0006] Preferably, the text generation image model based on the Tansformer model and the diffusion model includes: Text preprocessing: Get the input text and tokenize it, breaking it down into words or subwords; Text encoding: Use the encoder of the Transformer model to encode the input text after word segmentation and output the text vector; Text image generation: The text vector output by the encoder of the Transformer model is used as a conditional input to guide the diffusion model to generate a generated image related to the input text description; in each denoising process of the diffusion model, each layer of the generated image is affected by the conditional input, and finally a text image is output; Joint training: training a text encoder for text encoding and training a diffusion model for text image generation; performing joint training based on the text encoder and the diffusion model to construct the text generation image model; Model optimization: Define the first joint loss function and use it to evaluate the semantic consistency between the input text and the output text image; also introduce an additional discriminator to evaluate the quality of the text image and the consistency between the text image and the input text description.
[0007] Preferably, the image generation model based on the generative adversarial network model and the diffusion model comprises: Image preprocessing: Get the input image and perform image normalization and data augmentation on it; Generative adversarial network model: the preprocessed input image is input into the generator, and after several layers of convolution and deconvolution operations, the target image is generated according to the features of the input image; the discriminator is also used to extract the features of the input image through a deep convolutional neural network, and a probability value is output to judge the authenticity of the input image, and the judgment result is fed back to the generator; the generator and the discriminator are also trained through a binary cross entropy loss function; Diffusion model: The input image is gradually denoised through the diffusion process and is completely filled with noise. The diffusion process is reversed through the denoising process. Based on the characteristics of the input image, the noise of each step is predicted in multiple steps, and then all the noise is gradually removed to restore the original data. Model combination: combining the generator with the diffusion model, specifically: using the diffusion model as part of the generator, or introducing a denoising process at certain stages of the generator; combining the discriminator with the diffusion model, specifically: using the diffusion model to generate details of the input image to obtain the image generated by the diffusion model, and then providing additional discriminant signals through the discriminator to evaluate the quality of the image generated by the diffusion model; Model training and optimization: The denoising loss of the diffusion model and the adversarial loss of the generative adversarial network model together constitute a second joint loss function, and the model is jointly trained and optimized based on the second joint loss function to generate the image generation image model.
[0008] Preferably, the step of fusing the text-generated image model with the image-generated image model to generate the image-generated model comprises: A cross-modal embedding method is used to project text input data and image input data into a shared joint embedding space; Constructing a conditional diffusion model in the joint embedding space to generate a corresponding image according to the input text input data and image input data; Design a dual generator structure in the generative adversarial network model, one for generating images based on text input data, and the other for optimizing existing images; Establish a feedback loop mechanism so that the results of each round are used to optimize the results of the previous round; The recurrent feedback mechanism is incorporated into the joint embedding space to form the image generation model.
[0009] Preferably, the classifying the data input into the image generation model comprises: Get a mixed dataset containing text data and image data; Select and design a unified model framework; Performing model training on the unified model framework based on the mixed data set to generate a data classification model; The data input into the image generation model is classified by the data classification model.
[0010] Preferably, the selecting and designing of a unified model framework is specifically: Select the two-stream network structure in the neural network architecture as the unified model framework; Input text data into the text branch and extract text features through a recurrent neural network or Transformer; Input image data into the image branch and extract image features through the convolutional neural network; Performing feature fusion on the text feature and the image feature; The fused features are input into a fully connected layer for classification; The binary cross entropy loss function is used as the loss function and the optimizer is selected for model training.
[0011] Preferably, the fusing the text features and the image features comprises: Feature splicing: performing feature splicing on the text features and the image features; Weighted fusion: assign different weights to the concatenated features and sum them up; Attention Mechanism: A dynamic weight is assigned to each feature through the attention mechanism, and the weights of the two are dynamically adjusted by calculating the similarity between the text feature and the image feature.
[0012] Preferably, performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image comprises: Construct an image optimization model based on variational autoencoder and diffusion model; Performing image optimization on the first generated image and the second generated image based on the image optimization model to obtain a model optimized image; Wherein, constructing the image optimization model includes: Mapping an input image to a latent space using an encoder of the variational autoencoder; Introduce the diffusion model and define the comprehensive optimization loss function; Image optimization is performed in the latent space based on the comprehensive optimization loss function, and an optimized latent representation is generated using a diffusion process.
[0013] In a second aspect, an embodiment of the present application provides an AIGC-based image intelligent generation system, which applies the above-mentioned image intelligent generation method, including: Model building module: used to build image generation model; Input classification module: used to classify the data input into the image generation model into text input data and image input data; Data processing module: used for processing the text input data and the image input data to obtain text processing data and image processing data respectively; Image generation module: used for inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Image optimization module: used for performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Image content detection module: used to perform content detection on the model-optimized image, determine whether the image content complies with legal regulations and output the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
[0014] The beneficial effects of the present invention are as follows: the present invention effectively combines a variety of technologies such as the Tansformer model, the diffusion model, the generative adversarial network model, and the variational autoencoder to construct an image generation model, and uses the image generation model to generate images intelligently, which can generate images according to the input text information, and can also generate images according to the input image information, so it can comprehensively handle a variety of situations to achieve the comprehensiveness of image generation; and the present application can solve the problem of lack of realism caused by missing or unnatural image details through the diffusion model and the generative adversarial network model; improve the generalization ability of the model through the text generation image model and the image generation image model; and also realize cross-modal generation through the Tansformer model and the diffusion model; also judge whether the content of the generated image is compliant and true by detecting the image content, so as to ensure the credibility and standardization of the image content. The present application combines the above-mentioned multiple algorithms and models, improves the quality, interpretability and security of the generated image, and has high-quality visual effects, and can also be customized and optimized according to user needs, so as to achieve the comprehensiveness of intelligent image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] For better understanding and implementation, the technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0016] Figure 1 A flowchart of a method for intelligently generating an image based on AIGC provided in an embodiment of the present application; Figure 2 A flowchart of the steps of constructing an image generation model provided in an embodiment of the present application; Figure 3 A flowchart of the steps for generating an image generation model provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an AIGC-based intelligent image generation system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to further explain the technical means and effects taken by the present invention to achieve the predetermined invention purpose, exemplary embodiments will be described in detail here, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of methods and systems consistent with some aspects of the present application as detailed in the attached claims.
[0018] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this article refers to any or all possible combinations of one or more associated listed items.
[0019] The specific implementation methods, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0020] Example 1 See also Figure 1 The embodiment of the present application provides an AIGC-based image intelligent generation method, comprising the following steps: Build an image generation model; Classifying the data input into the image generation model into text input data and image input data; Processing the text input data and the image input data to obtain text processed data and image processed data respectively; Inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Performing content detection on the model-optimized image to determine whether the image content complies with legal regulations and outputting the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
[0021] Specifically, since the existing AIGC-based intelligent image generation may have many problems such as lack of details or unnaturalness resulting in lack of realism, the generalization ability of the model needs to be improved, cross-modal generation (such as text-generated images) is complex, and a large amount of data used for training is illegal or biased and the generated content does not comply with legal provisions and social norms, in order to avoid the generated content not complying with legal norms, lack of realism in the quality of generated images, insufficient generalization ability of the model itself and limitations of cross-modal generation, this application conducts further research and improvement on model construction, image generation (text-generated images and image-generated images) and image content detection, mainly including: first constructing an image generation model; then inputting the image generation model into the image generation model The data is classified into text input data and image input data; the text input data and the image input data are processed to obtain text processing data and image processing data respectively; the text processing data and the image processing data are then input into the image generation model to output a first generated image and a second generated image; the first generated image and the second generated image are then subjected to image optimization processing to obtain a model optimized image; finally, the model optimized image is subjected to content detection to determine whether the image content complies with legal regulations and output a final model generated image; wherein the first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
[0022] This application effectively combines multiple technologies such as Tansformer model, diffusion model, generative adversarial network model and variational autoencoder to construct an image generation model, and uses the image generation model to generate images intelligently. It can generate images based on input text information and image information, so it can comprehensively handle multiple situations to achieve comprehensive image generation; and this application can solve the problem of lack of realism caused by missing or unnatural image details through diffusion model and generative adversarial network model; improve the generalization ability of the model through text generation image model and image generation image model; and also realize cross-modal generation through Tansformer model and diffusion model; also judge whether the content of the generated image is compliant and true by detecting the image content, so as to ensure the credibility and standardization of the image content. This application combines the above-mentioned multiple algorithms and models to improve the quality, interpretability and security of the generated image, and has high-quality visual effects. At the same time, it can also be customized and optimized according to user needs (such as selecting text to generate images, or choosing to clarify blurred images), so as to achieve comprehensive image intelligent generation.
[0023] like Figure 2 As shown, in one embodiment provided in the present application, the step of constructing an image generation model includes: Construct a text-to-image model based on the Tansformer model and diffusion model; Construct an image generation model based on the generative adversarial network model and the diffusion model; The text generation image model and the image generation image model are fused to generate the image generation model.
[0024] Specifically, the Transformer model has excellent performance in sequence modeling, so it is very suitable for processing long-distance dependencies and contextual information. Especially when combined with text-generated image tasks, the Transformer model can better handle the association between natural language and visual information; while the diffusion model generates clear images from random noise by gradually denoising. Its core idea is to gradually transform a blurry image into a clear image through a multi-step reverse process. Therefore, the Tansformer model and the diffusion model are combined to construct a text-generated image model. First, the Transformer model is used to process the input text data and generate a matching latent space representation; then the diffusion model is used to generate images, which can generate detailed images while maintaining the semantics of the text. Since the diffusion model has a good ability to retain image details, it is very suitable for generating high-quality images.
[0025] Generative adversarial networks (GANs) are composed of two neural networks: the generator and the discriminator. The generator is responsible for creating images, while the discriminator evaluates the authenticity of the images. The generator and the discriminator are trained adversarially, which ultimately enables the generator to generate highly realistic images. The diffusion model can enhance the detail retention and diversity in the image generation process. Therefore, combining the advantages of the above two models can transform existing images into another form or style, generate higher quality images and convert between images, thereby improving the final generation effect.
[0026] Finally, this application combines the above-mentioned text-generated image model and image-generated image model to generate the final image generation model, fully combining the advantages of multiple models to improve the quality of the final image generation. And the final image generation model can not only distinguish text data to generate text images, but also distinguish image data to generate images, thereby comprehensively improving the accuracy and quality of the final generated image.
[0027] In an embodiment provided in the present application, the text generation image model is constructed based on the Tansformer model and the diffusion model, including: Text preprocessing: Get the input text and tokenize it, breaking it down into words or subwords; Text encoding: Use the encoder of the Transformer model to encode the input text after word segmentation and output a text vector. The Transformer network usually contains a multi-layer self-attention mechanism, which can capture long-distance dependencies and understand contextual information. The output of the encoder is a fixed-dimensional text vector (or vector sequence), which can effectively represent the input text semantics. This text vector will be used as a conditional input for the image generation process. Text image generation: The text vector output by the encoder of the Transformer model is used as a conditional input to guide the diffusion model to generate a generated image related to the input text description; in each denoising process of the diffusion model, each layer of the generated image is affected by the conditional input, and finally a text image is output; this conditional input can be combined with the noisy data through splicing, weighted summation or multi-level conditional generation strategy.
[0028] Joint training: training a text encoder for text encoding and training a diffusion model for text image generation; performing joint training based on the text encoder and the diffusion model to construct the text generation image model; Model optimization: Define the first joint loss function and use it to evaluate the semantic consistency between the input text and the output text image; also introduce an additional discriminator to evaluate the quality of the text image and the consistency between the text image and the input text description.
[0029] Specifically, this embodiment builds a text-generated image model based on the Tansformer model and the diffusion model. Its main process includes: text processing and encoding: using the Transformer model to convert text into potential representations (text vectors); image generation: using the conditional diffusion model to generate images that match the text description from noise; joint training and optimization: optimizing the model through joint training to ensure that the generated image has high consistency with the input text. This embodiment can effectively build a powerful text-generated image model by combining the powerful text understanding ability of the Transformer model with the high-quality image generation ability of the diffusion model.
[0030] It can be understood that the basic process of the diffusion model includes: Diffusion process (Forward Process): Starting from the real image, noise is gradually added through a series of steps until it eventually becomes pure noise. This process is a Markov chain, and each step adds a little noise, causing the image to gradually lose its original structure.
[0031] Denoising process (Reverse Process): It is the core of the diffusion model. It is an inverse process. The model learns how to restore the original image from the noise. In each step, the denoising network predicts how to remove the noise, thereby gradually restoring a clear image.
[0032] It should be noted that the input text in this embodiment is not the same as the above-mentioned text input data. The input text is the text information data required when building the image generation model; and the text input data refers to the text data that needs to be recognized and image generated in the input model after the image generation model is built. In other words, although both are text data, their functions are different. The former acts in the model training process, while the latter acts after the model is generated.
[0033] In an embodiment provided in the present application, the image generation image model is constructed based on the generative adversarial network model and the diffusion model, including: Image preprocessing: Get the input image and perform image normalization and data augmentation on it; perform augmentation operations on the training data, such as random cropping, rotation, flipping, etc., to improve the generalization ability of the model.
[0034] Generative adversarial network model: The preprocessed input image is input into the generator, and after several layers of convolution and deconvolution operations, the target image is generated according to the features of the input image; the discriminator is also used to extract the features of the input image through a deep convolutional neural network, and a probability value is output to judge the true or false of the input image, and the judgment result is fed back to the generator; the generator and the discriminator are also trained through a binary cross entropy loss function; it can be understood that the goal of the generator is to generate an image that matches the distribution of the target image; the optimization goal of the discriminator is to judge the true or false of the input image as accurately as possible, thereby forcing the generator to generate more realistic images.
[0035] Diffusion model: The input image is gradually denoised through the diffusion process and is completely filled with noise. This process usually consists of multiple time steps, and a small amount of noise is added at each step until the input image completely loses its original information. The denoising process is used to reverse the diffusion process, predict the noise at each step based on the characteristics of the input image, and then gradually remove all the noise to restore the original data. Model combination: combining the generator with the diffusion model, specifically: using the diffusion model as part of the generator, or introducing a denoising process at certain stages of the generator; combining the discriminator with the diffusion model, specifically: using the diffusion model to generate details of the input image to obtain the image generated by the diffusion model, and then using the discriminator to provide additional discriminant signals to evaluate the quality of the image generated by the diffusion model.
[0036] It can be understood that by using the diffusion model as part of the generator, or introducing a denoising process at certain stages of the generator, the images generated by the diffusion model can be fed into the discriminator of the generative adversarial network for further optimization. Specifically, the discriminator of the generative adversarial network can be applied to the preliminary images generated by the diffusion model to optimize the image quality and make it more realistic; on the other hand, the discriminator of the generative adversarial network can be used to evaluate the quality of the images generated by the diffusion model, making the generated images more realistic through adversarial training; by training a joint model, the diffusion model can be responsible for the details of the generated image, while the discriminator provides additional discriminative signals to improve the quality of the image.
[0037] Model training and optimization: The denoising loss of the diffusion model and the adversarial loss of the generative adversarial network model together constitute a second joint loss function, and the model is jointly trained and optimized based on the second joint loss function to generate the image generation image model; in the joint training stage, the second joint loss function is optimized, and the diffusion model and the generative adversarial network model are trained at the same time, so that the generated image has both high-quality visual effects and conforms to the semantic structure of the source image.
[0038] Specifically, this embodiment uses a generative adversarial network model and a diffusion model to construct an image generation model, which fully combines the advantages of the two algorithm models, can improve the quality of image generation, and the generated images have both high details and high realism; enhance the stability of training, reduce the common training instability and mode collapse problems of generative adversarial networks; enhance the detail performance of generated images, especially in high-frequency details (such as textures and edges); improve the efficiency of the training process and reduce complex optimization processes; provide powerful control capabilities to generate diverse and controllable images. In general, this combination can give full play to the respective advantages of generative adversarial networks and diffusion models, thereby improving the overall effect of image generation tasks.
[0039] like Figure 3 As shown, in one embodiment provided in the present application, the step of fusing the text generation image model and the image generation image model to generate the image generation model includes: A cross-modal embedding method is used to project text input data and image input data into a shared joint embedding space; Constructing a conditional diffusion model in the joint embedding space to generate a corresponding image according to the input text input data and image input data; Design a dual generator structure in the generative adversarial network model, one for generating images based on text input data, and the other for optimizing existing images; Establish a feedback loop mechanism so that the results of each round are used to optimize the results of the previous round; The recurrent feedback mechanism is incorporated into the joint embedding space to form the image generation model.
[0040] Specifically, the goal of this embodiment is to combine the advantages of the above two models to create a powerful and versatile model. The advantage of the text-generated image model is that it can generate visual content based on the input text, which is suitable for generating a variety of images; while the advantage of the image-generated image is that it converts the existing image into another form or style, can generate higher quality images and convert between images; therefore, when combining the above two models, there will be challenges such as model fusion and multimodal learning, specifically: Model fusion: The text-generated image model and the image-generated image model are different in structure and task objectives. How to effectively combine these two types of tasks is a core challenge; Multimodal learning: Effectively combine text information with image information, how to make the image generation process understand the text description, and seamlessly integrate the conversion between images (such as style transfer, image restoration, etc.). Therefore, in response to the above problems, this embodiment first uses cross-modal embedding to project text and images into a shared latent space to ensure the matching between text and image information. In this way, both the text-generated image model and the image-generated image model can be optimized in the same latent space; secondly, a conditional diffusion model is constructed, which can generate corresponding images based on the input text and image conditions. For example, by inputting a combination of text and image, a new image is generated in which both the image content and style are changed. Then in the GAN part, a dual generator structure is designed, one for generating text-based images and the other for optimizing existing images. The discriminator of the generator can optimize the generation process by judging the authenticity of the image and the degree of match between the text description and the image. This method ensures that the generated image is not only visually realistic, but also consistent with the input text description. Finally, a loop feedback mechanism is established. By using the image-text-image loop generation method, each round of generation can optimize the results of the previous round. The final generated image not only conforms to the text description, but is also highly optimized in terms of style and details.
[0041] Therefore, by combining the two tasks of text-generated image and image-generated image in the above way, we can not only generate high-quality images based on text, but also process and generate various image transformation tasks (such as style transfer, deblurring, super-resolution, etc.). The advantages of this model are: Text-driven image generation: Able to generate images with high consistency based on text descriptions.
[0042] Image-to-image generation capability: Ability to generate images with optimized style, quality, and details.
[0043] Multimodal optimization: By sharing latent spaces and conditional generation strategies, images and text can be seamlessly connected, thereby improving the quality and diversity of image generation.
[0044] This combination can build a currently very advanced and powerful comprehensive image generation model, achieving the best results in the field of image generation.
[0045] In an embodiment provided in the present application, the classifying the data input into the image generation model includes: Get a mixed dataset containing text data and image data; Select and design a unified model framework; Performing model training on the unified model framework based on the mixed data set to generate a data classification model; The data input into the image generation model is classified by the data classification model.
[0046] In an embodiment provided in the present application, the selection and design of a unified model framework is specifically as follows: Select the two-stream network structure in the neural network architecture as the unified model framework; Input text data into the text branch and extract text features through a recurrent neural network or Transformer; Input image data into the image branch and extract image features through the convolutional neural network; Performing feature fusion on the text feature and the image feature; The fused features are input into a fully connected layer for classification; The binary cross entropy loss function is used as the loss function and the optimizer is selected for model training.
[0047] Specifically, this embodiment designs a unified model framework that can process text and image data at the same time, and process them through different branches or channels, and finally fuse them. In this way, the generated data classification model can simultaneously process and fuse features from text and images to achieve multimodal classification tasks.
[0048] In an embodiment provided in the present application, the step of fusing the text feature with the image feature includes: Feature splicing: performing feature splicing on the text features and the image features; Weighted fusion: assign different weights to the concatenated features and sum them up; Attention Mechanism: A dynamic weight is assigned to each feature through the attention mechanism, and the weights of the two are dynamically adjusted by calculating the similarity between the text feature and the image feature.
[0049] Specifically, this embodiment performs feature concatenation and weighted fusion, and establishes an attention mechanism, and can simultaneously utilize these three methods to process and optimize features of different modalities at different levels, enabling the model to better capture the relationship between modalities, while flexibly adjusting the contribution of each modality, thereby improving the performance of the multimodal model.
[0050] In an embodiment provided in the present application, performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image includes: Construct an image optimization model based on variational autoencoder and diffusion model; Performing image optimization on the first generated image and the second generated image based on the image optimization model to obtain a model optimized image; Wherein, constructing the image optimization model includes: Mapping an input image to a latent space using an encoder of the variational autoencoder; Introduce the diffusion model and define the comprehensive optimization loss function; Performing image optimization in the latent space based on the comprehensive optimization loss function, and generating an optimized latent representation using a diffusion process; The comprehensive optimization loss function is expressed as: ; in, is the VAE loss of the variational autoencoder; is the denoising loss of the diffusion model; The loss weight to balance the VAE loss and denoising loss.
[0051] Specifically, this embodiment combines the advantages of the Variational Autoencoder (VAE) and the diffusion model, which can improve the quality of image generation and make the generated image details more realistic and refined; it can also enhance the flexibility and multi-task capability of the model, and can simultaneously optimize multiple aspects of the image; as well as improve the diversity and stability of generation, avoid mode collapse, and increase the diversity of image generation; significantly improve image restoration and super-resolution tasks, and effectively restore image details; enhance control and interpretability, and the operability of the latent space makes image generation more controllable; balance the generation speed and quality, achieve a better balance in generation effect, and have a higher denoising ability.
[0052] Example 2 See also Figure 4 The embodiment of the present application provides an AIGC-based image intelligent generation system, which applies the above-mentioned AIGC-based image intelligent generation method, including: Model building module: used to build image generation model; Input classification module: used to classify the data input into the image generation model into text input data and image input data; Data processing module: used for processing the text input data and the image input data to obtain text processing data and image processing data respectively; Image generation module: used for inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Image optimization module: used for performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Image content detection module: used to perform content detection on the model-optimized image, determine whether the image content complies with legal regulations and output the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
[0053] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0054] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0055] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0056] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An AIGC-based image intelligent generation method, characterized by: The steps include: Build an image generation model; Classifying the data input into the image generation model into text input data and image input data; Processing the text input data and the image input data to obtain text processed data and image processed data respectively; Inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Performing content detection on the model-optimized image to determine whether the image content complies with legal regulations and outputting the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
2. The method for intelligently generating images based on AIGC according to claim 1, characterized in that: The constructing of the image generation model comprises: Construct a text-to-image model based on the Tansformer model and diffusion model; Construct an image generation model based on the generative adversarial network model and the diffusion model; The text generation image model and the image generation image model are fused to generate the image generation model.
3. The method for intelligently generating images based on AIGC according to claim 2, characterized in that: The text generation image model based on the Tansformer model and the diffusion model includes: Text preprocessing: Get the input text and tokenize it, breaking it down into words or subwords; Text encoding: Use the encoder of the Transformer model to encode the input text after word segmentation and output the text vector; Text image generation: The text vector output by the encoder of the Transformer model is used as a conditional input to guide the diffusion model to generate a generated image related to the input text description; in each denoising process of the diffusion model, each layer of the generated image is affected by the conditional input, and finally a text image is output; Joint training: training a text encoder for text encoding and training a diffusion model for text image generation; performing joint training based on the text encoder and the diffusion model to construct the text generation image model; Model optimization: Define the first joint loss function and use it to evaluate the semantic consistency between the input text and the output text image; also introduce an additional discriminator to evaluate the quality of the text image and the consistency between the text image and the input text description.
4. The method for intelligently generating images based on AIGC according to claim 2, characterized in that: The image generation image model based on the generative adversarial network model and the diffusion model includes: Image preprocessing: Get the input image and perform image normalization and data augmentation on it; Generative adversarial network model: the preprocessed input image is input into the generator, and after several layers of convolution and deconvolution operations, the target image is generated according to the features of the input image; the discriminator is also used to extract the features of the input image through a deep convolutional neural network, and a probability value is output to judge the authenticity of the input image, and the judgment result is fed back to the generator; the generator and the discriminator are also trained through a binary cross entropy loss function; Diffusion model: The input image is gradually denoised through the diffusion process and is completely filled with noise. The diffusion process is reversed through the denoising process. Based on the characteristics of the input image, the noise of each step is predicted in multiple steps, and then all the noise is gradually removed to restore the original data. Model combination: combining the generator with the diffusion model, specifically: using the diffusion model as part of the generator, or introducing a denoising process at certain stages of the generator; combining the discriminator with the diffusion model, specifically: using the diffusion model to generate details of the input image to obtain the image generated by the diffusion model, and then providing additional discriminant signals through the discriminator to evaluate the quality of the image generated by the diffusion model; Model training and optimization: The denoising loss of the diffusion model and the adversarial loss of the generative adversarial network model together constitute a second joint loss function, and the model is jointly trained and optimized based on the second joint loss function to generate the image generation image model.
5. The method for intelligently generating images based on AIGC according to claim 2, characterized in that: The step of fusing the text-generated image model with the image-generated image model to generate the image-generated model includes: A cross-modal embedding method is used to project text input data and image input data into a shared joint embedding space; Constructing a conditional diffusion model in the joint embedding space to generate a corresponding image according to the input text input data and image input data; Design a dual generator structure in the generative adversarial network model, one for generating images based on text input data, and the other for optimizing existing images; Establish a feedback loop mechanism so that the results of each round are used to optimize the results of the previous round; The recurrent feedback mechanism is incorporated into the joint embedding space to form the image generation model.
6. The method for intelligently generating images based on AIGC according to claim 1, characterized in that: The classifying the data input into the image generation model comprises: Get a mixed dataset containing text data and image data; Select and design a unified model framework; Performing model training on the unified model framework based on the mixed data set to generate a data classification model; The data input into the image generation model is classified by the data classification model.
7. The method for intelligently generating images based on AIGC according to claim 6, characterized in that: The selection and design of a unified model framework is as follows: Select the two-stream network structure in the neural network architecture as the unified model framework; Input text data into the text branch and extract text features through a recurrent neural network or Transformer; Input image data into the image branch and extract image features through the convolutional neural network; Performing feature fusion on the text features and the image features; The fused features are input into a fully connected layer for classification; The binary cross entropy loss function is used as the loss function and the optimizer is selected for model training.
8. The method for intelligently generating images based on AIGC according to claim 7, characterized in that: The step of fusing the text feature with the image feature includes: Feature splicing: performing feature splicing on the text features and the image features; Weighted fusion: assign different weights to the concatenated features and sum them up; Attention Mechanism: A dynamic weight is assigned to each feature through the attention mechanism, and the weights of the two are dynamically adjusted by calculating the similarity between the text feature and the image feature.
9. The method for intelligently generating images based on AIGC according to claim 1, characterized in that: The performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image comprises: Construct an image optimization model based on variational autoencoder and diffusion model; Performing image optimization on the first generated image and the second generated image based on the image optimization model to obtain a model optimized image; Wherein, constructing the image optimization model includes: Mapping an input image to a latent space using an encoder of the variational autoencoder; Introduce the diffusion model and define the comprehensive optimization loss function; Image optimization is performed in the latent space based on the comprehensive optimization loss function, and an optimized latent representation is generated using a diffusion process.
10. An AIGC-based image intelligent generation system, applying the image intelligent generation method according to any one of claims 1 to 9, characterized in that: include Model building module: used to build image generation model; Input classification module: used to classify the data input into the image generation model into text input data and image input data; Data processing module: used for processing the text input data and the image input data to obtain text processing data and image processing data respectively; Image generation module: used for inputting the text processing data and the image processing data into the image generation model to output a first generated image and a second generated image; Image optimization module: used for performing image optimization processing on the first generated image and the second generated image to obtain a model optimized image; Image content detection module: used to perform content detection on the model-optimized image, determine whether the image content complies with legal regulations and output the final model-generated image; The first generated image is a text generated image generated based on the text processing data; the second generated image is an image generated image generated based on the image processing data.
Citation Information
Patent Citations
Image processing method and apparatus, computer device, storage medium and product
WO2025001894A1
Cited By
Image generation training method for AIGC large model
CN121437670A