Image generation method and system based on adaptive image coding
By building adaptive image coding module and data enhancement technology, multi-size high-resolution learning of the image generation model is solved, and the performance and generalization problems of existing models are solved when generating high-definition images, achieving efficient generation of high-definition images.
Patent Information
- Application Number
- CN202510362660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing image generation model generates high-definition 2k or more and the detailed picture of the picture content, the calculation complexity and computing resources are limited by training samples, model training, resulting in reduced performance and generalization, and high post-processing costs.
The image data set is constructed for data enhancement, and the training data set is segmented through the adaptive image encoding module, and combined with the diffusion model for training. The generated model has the ability to learn multi-size high-resolution, and generates high-definition and higher-resolution target images.
It improves the generalization ability of the model, and can generate higher-definition and higher resolution target images without relying on post-processing, reducing image information loss and improving generation effect.
Smart Images

Figure CN120298520A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image generation method and system based on adaptive image coding. Background Art
[0002] The AIGC technology can not only greatly improve the content production efficiency, but also show unique creativity and high-quality output in some fields. The SD model is also an image generation model based on LDM introduced in recent years. This model combines the advantages of DDPM and VQ-VAE, and adds a conditional image generation module, which generates images conditionally according to the user input text description, thus making significant progress in the quality, speed and cost of image generation. However, due to the limitations of training samples, the computational complexity of model training, and computing power resources, a large number of large-size data in the training data are discarded or the image size is scaled. The loss of these high-definition picture information is likely to lead to a decrease in model performance and generalization, and the effect of generating pictures with a resolution higher than 2k and more realistic picture content details is not ideal. In today's era of constantly pursuing picture quality, there are also methods such as high-definition restoration and other post-processing, but they are all limited by the GPU video memory and cannot output pictures stably. At the same time, the time-consuming cost of post-processing high-definition pictures is also longer. Summary of the Invention
[0003] To solve the above technical problems, the purpose of the present invention is to provide an image generation method and system based on adaptive image coding that has the ability to generate pictures with high-definition resolution.
[0004] To achieve the above object, one aspect of the embodiments of the present application proposes an image generation method based on adaptive image coding, including the following steps:
[0005] Construct an image data set, perform data augmentation on the image data set to obtain a first training data set;
[0006] Construct an adaptive image coding module, and perform image segmentation on the first training data set through the adaptive image coding module to obtain a second training data set;
[0007] Train a preset diffusion model according to the second training data set to obtain an image generation model;
[0008] Obtain user requirement information, input the user requirement information into the image generation model to obtain a target image.
[0009] In some embodiments, the constructing of the image data set specifically includes:
[0010] Collect multiple types of first training images, and screen the first training images according to a preset image screening condition to obtain second training images;
[0011] Filter the second training images to obtain third training images;
[0012] Annotate the third training images to obtain fourth training images;
[0013] Divide the fourth training images into a training set, a validation set, and a test set to obtain the image dataset.
[0014] In some embodiments, the image splitting of the first training dataset by the adaptive image encoding module to obtain a second training dataset specifically includes:
[0015] Perform sub - graph splitting on the first training dataset to obtain sub - graph position encoding information and sub - graph size information;
[0016] Perform size scaling on the first training dataset to obtain original image size information;
[0017] Concatenate the sub - graph position encoding information, the sub - graph size information, and the original image size information to obtain the second training dataset.
[0018] In some embodiments, the first training dataset includes multiple original images, and the performing sub - graph splitting on the first training dataset to obtain sub - graph position encoding information and sub - graph size information specifically includes:
[0019] Determine the target image size and the target aspect ratio, and calculate the original image size of the original image;
[0020] Determine the initial splitting times according to the original image size and the target image size, and then generate a candidate splitting times set according to the initial splitting times;
[0021] Determine the splitting rule, and according to the candidate splitting times set and the splitting rule, obtain a list of candidate splitting schemes, where the list of candidate splitting schemes includes multiple sub - graph aspect ratios;
[0022] Calculate the proximity ratio of each sub - graph aspect ratio to the target aspect ratio, and then determine the best sub - graph aspect ratio according to the proximity ratio;
[0023] According to the best sub - graph aspect ratio and the original image size, split each original image to obtain the sub - graph position encoding information and the sub - graph size information.
[0024] In some embodiments, the step of segmenting each of the original images according to the optimal sub-image aspect ratio and the original image size to obtain the sub-image position encoding information and the sub-image size information specifically includes:
[0025] Calculating the sub-image size information according to the original image size and the optimal sub-image aspect ratio;
[0026] Segmenting the original image according to the sub-image size information to obtain a plurality of sub-images;
[0027] Calculating the position coordinates of each sub-image in the original image to obtain the sub-image position encoding information.
[0028] In some embodiments, the step of splicing the sub-image position encoding information, the sub-image size information, and the original image size information to obtain the second training data set specifically includes:
[0029] Converting the sub-image position encoding information, the sub-image size information, and the original image size information into Fourier encodings;
[0030] Splicing the Fourier encodings to obtain the second training data set.
[0031] In some embodiments, the diffusion model includes a U-Net module. The step of training a preset diffusion model according to the second training data set to obtain an image generation model specifically includes:
[0032] Setting training parameters and freezing the weights of the U-Net module;
[0033] Inputting the second training data set into the diffusion model and performing fine-tuning training on the diffusion model according to the training parameters;
[0034] Evaluating the performance of the fine-tuned diffusion model according to the Fréchet inception distance and the CLIP score to obtain a performance evaluation result;
[0035] Updating the parameters of the fine-tuned diffusion model according to the performance evaluation result to obtain the image generation model.
[0036] To achieve the above object, another aspect of the embodiments of the present application proposes an image generation system based on adaptive image encoding, including:
[0037] A data augmentation module for constructing an image data set and performing data augmentation on the image data set to obtain a first training data set;
[0038] An adaptive processing module, configured to construct an adaptive image encoding module, and perform image segmentation on the first training dataset through the adaptive image encoding module to obtain a second training dataset;
[0039] A model training module, configured to train a preset diffusion model according to the second training dataset to obtain the image generation model;
[0040] An image generation module, configured to obtain user requirement information, input the user requirement information into the image generation model, and obtain a target image.
[0041] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the image generation method based on adaptive image encoding as described above.
[0042] To achieve the above object, on the other hand, an embodiment of the present application provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the image generation method based on adaptive image encoding as described above.
[0043] The beneficial effects of the present invention are as follows: For the image generation method and system based on adaptive image encoding of the present invention, first, an image dataset is constructed, and data augmentation is performed on the image dataset to obtain a first training dataset. Then, an adaptive image encoding module is constructed, and image segmentation is performed on the first training dataset through the adaptive image encoding module to obtain a second training dataset. Furthermore, a preset diffusion model is trained according to the second training dataset to obtain an image generation model. Finally, user requirement information is input into the image generation model to obtain a target image. In the process of training the image generation model of the present invention, a preposed adaptive image encoding module is added to perform image segmentation on large-size pictures in the training dataset, which can enable the image generation model to perform multi-size and high-resolution learning, improve the generalization ability of the model, and through this image generation model, a target image that is clearer and has a higher resolution can be generated without relying on post-processing. Description of the Drawings
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a flowchart of steps of an image generation method based on adaptive image coding provided by an embodiment of the present invention;
[0046] Figure 2 It is a flowchart of steps of image dataset construction provided by an embodiment of the present invention;
[0047] Figure 3 It is a flowchart of steps of processing by an adaptive image coding module provided by an embodiment of the present invention;
[0048] Figure 4 It is a flowchart of steps of step S1021 provided by an embodiment of the present invention;
[0049] Figure 5 It is an example diagram of a candidate segmentation scheme provided by an embodiment of the present invention;
[0050] Figure 6 It is a flowchart of steps of step S10215 provided by an embodiment of the present invention;
[0051] Figure 7 It is a flowchart of steps of step S1023 provided by an embodiment of the present invention;
[0052] Figure 8 It is a flowchart of steps of step S103 provided by an embodiment of the present invention;
[0053] Figure 9 It is a schematic flowchart of an image generation method based on adaptive image coding provided by an embodiment of the present invention;
[0054] Figure 10 It is a schematic structural diagram of an image generation system based on adaptive image coding provided by an embodiment of the present invention;
[0055] Figure 11 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not used to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of devices and methods that are consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0057] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information. Depending on the context, words such as "if" and "when" used herein may be interpreted as "when...", "while...", or "in response to determining".
[0058] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0059] Before elaborating in detail on the embodiments of this application, first, some nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0060] AIGC (AI Generated Content) is a new content production method that uses generative artificial intelligence technology to automatically create content such as text, images, and videos. It relies on technologies such as deep learning, generative models, natural language processing, and computer vision, and has advantages such as high efficiency, diversity, and innovation. It is widely used in fields such as content creation, commercial marketing, education, entertainment, and scientific research.
[0061] Stable Diffusion (SD) is a generative artificial intelligence model based on the diffusion model, mainly used for generating high-quality images. It generates realistic images step by step from random noise by simulating the diffusion process and is widely used in fields such as art creation, design, and game development.
[0062] LDM (Latent Diffusion Model) is a deep generative model that operates on the low-dimensional latent space of a pre-trained autoencoder rather than directly on the pixel space. This makes it less costly, capable of rapid sampling and efficient training, and reduces computational resource consumption.
[0063] VQ-VAE (Vector Quantized Variational AutoEncoder) is a deep learning model that combines the ideas of vector quantization and variational autoencoder to learn effective representations of data and generate high-quality samples.
[0064] U-Net is a deep learning-based convolutional neural network mainly used for image segmentation tasks, especially performing well in the field of biomedical image segmentation. It consists of two parts: an encoder (downsampling path) and a decoder (upsampling path), shaped like a U. The core lies in its encoder-decoder architecture and skip connections, which enable the network to accurately perform pixel-level classification and segmentation.
[0065] AIGC technology can not only significantly improve the content production efficiency but also demonstrate unique creativity and high-quality output in certain fields. The SD model is also an image generation model based on LDM introduced in recent years. This model combines the advantages of DDPM and VQ-VAE and adds a conditional image generation module to generate images conditionally according to the text descriptions input by users, thus making significant progress in the quality, speed, and cost of image generation. However, due to limitations in training samples, the computational complexity of model training, and computing power resources, a large number of large-sized data in the training data are discarded or resized, and the loss of information in these high-definition pictures is likely to lead to a reduction in model performance and generalization ability, resulting in unsatisfactory results for generating pictures with a resolution higher than 2k and more realistic details in the picture content. In today's era of constantly pursuing image quality, there are also methods such as high-definition restoration and other post-processing, but they are all restricted by the GPU video memory and cannot stably generate images. At the same time, post-processing high-definition pictures also takes longer in terms of time and cost.
[0066] To this end, an embodiment of the present invention proposes an image generation method based on adaptive image coding. First, an image data set is constructed, and data augmentation is performed on the image data set to obtain a first training data set. Then, an adaptive image coding module is constructed, and the first training data set is segmented by the adaptive image coding module to obtain a second training data set. Furthermore, a preset diffusion model is trained according to the second training data set to obtain an image generation model. Finally, user requirement information is input into the image generation model to obtain a target image. In the process of training the image generation model, the present invention adds a pre - adaptive image coding module to segment large - size pictures in the training data set, enabling the image generation model to perform multi - size and high - resolution learning, improving the generalization ability of the model. Through this image generation model, a target image with higher definition and higher resolution can be generated without relying on post - processing. This image generation method can be applied to scenarios such as art creation, advertising design, and educational material generation, but is not limited thereto.
[0067] Referring to Figure 1 , Figure 1 FIG. is a flowchart of steps of an image generation method based on adaptive image coding provided by an embodiment of the present invention. An embodiment of the present invention proposes an image generation method based on adaptive image coding, and this method includes steps S101 to S104:
[0068] S101. Construct an image data set, perform data augmentation on the image data set, and obtain a first training data set;
[0069] Specifically, construct a rich image data set, which includes an open - source picture data set, an Internet - crawled data set, etc. Considering that the effect of the model - generated pictures is determined by the data set, a data augmentation method is adopted to expand the training set and increase the diversity of data categories. For example, some pictures are enhanced by offline high - definition magnification to increase the proportion of large - size samples. Some pictures first adjust the image size so that the shortest side matches the target size, and then randomly crop or centrally crop the image along the longer side, but ensure that the main part of the picture is not missing after cropping.
[0070] It should be noted that the image data set constructed in the embodiment of the present invention forms a first training data set after data annotation and augmentation, which not only contains rich sample features but also takes into account the stability and generalization ability of model training, contributing to subsequent model training.
[0071] Referring to Figure 2 , Figure 2 FIG. is a flowchart of steps for constructing an image data set provided by an embodiment of the present invention. Further, as an optional implementation manner, the step of constructing an image data set can be further divided into the following steps S1011 to S1014:
[0072] S1011. Collect multiple types of first training images, and screen the first training images according to preset image screening conditions to obtain second training images;
[0073] S1012. Filter the second training images to obtain third training images;
[0074] S1013. Annotate the third training images to obtain fourth training images;
[0075] S1014. Divide the fourth training images into a training set, a validation set, and a test set to obtain an image dataset.
[0076] In some alternative embodiments, first collect a large number of images with different themes, painting styles, and concepts through open-source image datasets, Internet crawled datasets, etc., and screen, filter, and annotate the first training images from different sources. The size of the data is preferably greater than 300K, and the data size needs to be above 512x512 pixels; screen out data with low resolution, poor quality (such as pictures with a resolution of 768*768 < 100kb), damage, and data irrelevant to the task objective; remove contamination features such as watermarks and interfering text that may be contained in the data. Data annotation can be divided into automatic annotation and manual annotation. Automatic annotation mainly relies on models that can generate labels for pictures, such as BLIP (img2caption) and Waifu Diffusion 1.4 (img2tag), while manual annotation relies on annotators. Annotate the text information in the third training images as the prompt words for diffusion model training to obtain fourth training images. Finally, count the total number of images N, and then randomly divide the fourth training images into a training set, a validation set, and a test set according to the ratio of 8:1:1 to obtain an image dataset.
[0077] Furthermore, considering the impact of training data on the model learning effect, adopt a data augmentation strategy to expand the training set samples, and perform Image Augmentation processing on each picture repeatedly: ① Add random angle rotation to the picture to simulate the shooting angle and change the target position to achieve the change of the target movement position; ② Perform multi-size augmentation on the data, such as performing cropping and scaling operations with sizes of 1:1, 1:2, 2:1, 1:3, 3:4, 4:3, 9:16, and 16:9, but retain the main features of the picture (such as faces, buildings, etc.) during multi-size augmentation; ③ Perform high-definition magnification processing on some pictures to increase the proportion of large-size samples in the samples; ④ Increase / decrease the hue, saturation, and brightness values for color transformation to restore the impact of lighting conditions, thereby improving the quality and diversity of the training set pictures and being beneficial to the improvement of the model generation effect.
[0078] S102. Construct an adaptive image coding module, and perform image segmentation on the first training dataset through the adaptive image coding module to obtain a second training dataset;
[0079] Specifically, an adaptive image coding module is added based on the open-source SD1.5 model, enabling the model to have training strategies for more images with higher resolutions and larger sizes. Meanwhile, the image information is retained without loss, improving the overall quality of the images and the control of local details.
[0080] Refer to Figure 3 , Figure 3 FIG. is a flowchart of steps for processing by the adaptive image coding module provided in an embodiment of the present invention. Further, as an optional implementation manner, the step of performing image segmentation on the first training dataset through the adaptive image coding module to obtain a second training dataset can be specifically further divided into the following steps S1021 to S1023:
[0081] S1021. Perform sub-image segmentation on the first training dataset to obtain sub-image position coding information and sub-image size information;
[0082] Specifically, use the adaptive image coding module to preprocess the first training dataset after data augmentation, and perform sub-image segmentation on images with a resolution above 1024 to obtain sub-image position coding information and sub-image size information.
[0083] Refer to Figure 4 , Figure 4 FIG. is a flowchart of steps for step S1021 provided in an embodiment of the present invention. Further, as an optional implementation manner, the first training dataset includes multiple original images. The step of performing sub-image segmentation on the first training dataset to obtain sub-image position coding information and sub-image size information can be specifically further divided into the following steps S10211 to S10215:
[0084] S10211. Determine the target image size and the target aspect ratio, and calculate the original image size of the original image;
[0085] S10212. Determine the initial segmentation times according to the original image size and the target image size, and then generate a candidate segmentation times set according to the initial segmentation times;
[0086] In some alternative embodiments, if the target image size is determined to be width \(t_w\) and height \(t_h\), then the target aspect ratio is \(t_w / t_h\). Randomly select an original image \(x\) from the first training dataset, and calculate the original image size (width \(w\) and height \(h\)) of this original image. According to the ratio of the original image size to the target image size such as 512*512, if the original image size is greater than 512*512, then determine the initial number of cuts \(N\) through the following formula, round the calculation result, and set \([N - 1, N, N + 1]\) as the candidate set of the number of cuts; if the original image size is less than 512*512, then do not perform cutting.
[0087] \(N=\min(h\times w / t_h\times t_w,\max\_num)\)
[0088] Where \(N\) represents the initial number of cuts, \(w\) and \(h\) represent the width and height of the original image size, \(t_w\) and \(t_h\) represent the width and height of the target image size, and \(\max\_num\) represents the maximum number of cuts.
[0089] S10213. Determine the cutting rule. According to the candidate set of the number of cuts and the cutting rule, obtain a list of candidate cutting schemes. The list of candidate cutting schemes includes multiple sub - graph aspect ratios;
[0090] In some alternative embodiments, determine the cutting rule as: the ratio of the width * height of the sub - graph after cutting = the candidate number of cuts. Traverse all candidate numbers of cuts to obtain a list of candidate cutting schemes. Exemplarily, if the candidate set of the number of cuts is \([6,7,8]\), then the obtained list of candidate cutting schemes is \([[1,6],[2,3],[3,2],[6,1],[1,7],[7,1],[1,8],[2,4],[4,2],[8,1]]\), as Figure 5 shown.
[0091] S10214. Calculate the proximity ratio of each sub - graph aspect ratio to the target aspect ratio, and then determine the best sub - graph aspect ratio according to the proximity ratio;
[0092] In some alternative embodiments, traverse all candidate cutting schemes, find the closest ratio of the sub - graph aspect ratio after cutting to the target aspect ratio through the following calculation process, and determine the best sub - graph aspect ratio \([a,b]\) according to the closest ratio. For example, select the 2x3 cutting scheme.
[0093] \(cur_w = w / x\)
[0094] \(cur_h = h / y\)
[0095] \(\min(\vert\log(cur_w / cur_h)-\log(t_w / t_h)\vert)\)
[0096] Among them, w and h represent the width and height of the original image size, t_w and t_h represent the width and height of the target image size, and [x, y] represents the candidate segmentation scheme.
[0097] S10215. Segment each original image according to the optimal sub-image aspect ratio and the original image size to obtain sub-image position encoding information and sub-image size information.
[0098] Refer to Figure 6 , Figure 6 FIG. is a flowchart of a step of step S10215 provided by an embodiment of the present invention. Further as an optional implementation manner, the step of segmenting each original image according to the optimal sub-image aspect ratio and the original image size to obtain sub-image position encoding information and sub-image size information can be specifically further divided into the following steps A101 to A103:
[0099] A101. Calculate sub-image size information according to the original image size and the optimal sub-image aspect ratio;
[0100] A102. Segment the original image according to the sub-image size information to obtain a plurality of sub-images;
[0101] Specifically, according to the original image size (width w, height h) and the optimal sub-image aspect ratio [a, b], the sub-image size information corresponding to each sub-image is width * height = w / a * h / b. Then, use the cropping function to segment the original image according to the sub-image size information to obtain a plurality of sub-images.
[0102] A103. Calculate the position coordinates of each sub-image in the original image to obtain sub-image position encoding information.
[0103] Specifically, for each sub-image, calculate its position coordinates (x1, y1) in the original image through 2D position encoding information. The index of the sub-image is (i, j), where i ∈ [0, a - 1] and j ∈ [0, b - 1]. Then the center point coordinates of the sub-image (i.e., the sub-image position encoding information) are:
[0104] x1 = (i + 0.5)w / a
[0105] y1 = (j + 0.5)h / b
[0106] S1022. Scale the size of the first training data set to obtain original image size information;
[0107] S1023. Concatenate the sub-image position encoding information, the sub-image size information, and the original image size information to obtain a second training data set.
[0108] Specifically, according to the sub - figure size, scale the original image size and place a copy separately in the sub - figure list. In the second training dataset, there are both the scaled - down version of the original image size and each sub - figure segmented from the image. Retain the sub - figure position encoding information, sub - figure size information, original image size information, etc. Together, the segmented sub - figures and the scaled - down full image of the original figure are put into the second training dataset for the model to perform multi - size and higher - resolution learning.
[0109] Refer to Figure 7 , Figure 7 FIG. is a flowchart of a step of step S1023 provided by an embodiment of the present invention. Further, as an optional implementation manner, the step of splicing the sub - figure position encoding information, sub - figure size information, and original image size information to obtain the second training dataset can be further divided into the following steps S10231 and S10232:
[0110] S10231: Convert the sub - figure position encoding information, sub - figure size information, and original image size information into Fourier encoding;
[0111] S10232: Splice each Fourier encoding to obtain the second training dataset.
[0112] Specifically, the 2D position encoding information indicates the local position of each sub - figure in the original figure. After the coordinates are Fourier - encoded, they are added to the Time Embedding and used as additional conditional embeddings together with the original image size into the U - Net model. Thus, during the training process, the model can learn that a certain picture has been segmented, enhancing the understanding of "image segmentation" and the original resolution information of the image. Therefore, during the inference and generation stage, it can better adapt to the generation of images of different sizes and improve the generation effect of image detail information without generating noise artifacts.
[0113] S103: Train a preset diffusion model according to the second training dataset to obtain an image generation model;
[0114] In some optional embodiments, the diffusion model can adopt the open - source SD1.5 model for adaptive fine - tuning training of the image module. Freeze the weights of the existing U - Net model and perform full - scale parameter fine - tuning. The training samples are processed by the pre - placed adaptive image encoding module to enable the model to perform multi - size and higher - resolution learning. Use an independent test set to perform model evaluation, and select the Frechet inception distance and CLIP score metrics as evaluation metrics to measure the model performance; secondly, analyze the test results to determine whether the new model performs well in generating resolutions above the original 512 size, whether the details of the generated pictures are rich, and in which aspects there is still room for improvement.
[0115] Refer to Figure 8 ,Figure 8 This is a flowchart of step S103 provided by an embodiment of the present invention. As a further optional implementation, the diffusion model includes a U-Net module. The step of training a preset diffusion model according to a second training dataset to obtain an image generation model can be further divided into the following steps S1031 to S1034:
[0116] S1031. Set training parameters and freeze the weights of the U-Net module;
[0117] S1032. Input the second training dataset into the diffusion model and perform fine-tuning training on the diffusion model according to the training parameters;
[0118] Specifically, in the full-parameter fine-tuning training, the selection of the base model needs to choose an SD model with a generation ability distribution approximate to the training data distribution as the training base model (for example, when training a second-generation character dataset, an SD model with strong second-generation image generation ability can be selected). During the fine-tuning training of the SD model, continuous expansion and optimization learning are carried out on many abilities and concepts of the original base model, so as to obtain a comprehensive ability of the base model and the dataset distribution. Using the xformers acceleration library can accelerate the training of the SD model by about 2 times, because it can reduce the training video memory occupancy by 2 times, so that the Batch Size number can be increased. Then set the training parameters. The setting of the learning rate in the training process of the SD model is very sensitive. If the learning rate is set too large, it is very likely that the training of the SD model will go out of control and generate very poor pictures during forward inference; if the learning rate is set too small, the model may not be able to jump out of the minimum point. Therefore, for the total Batch Size (single-card Batch Size x number of GPUs) less than 10, the learning rate can be set to 2e-6; for the total Batch Size greater than 10 and less than 100, the learning rate can be set to 1e-7. Other training parameter configurations can be configured in the training framework config, such as: lr_scheduler, optimizer_type, lr_warmup_steps, etc. Package the training script in the accelerate library, start the configured training environment, and at the same time pass the just-configured config_file and sample_prompt parameters into the training script. Run the training script to perform full-parameter fine-tuning training on the SD model.
[0119] S1033. Evaluate the performance of the fine-tuned diffusion model according to the Fréchet inception distance and the CLIP score to obtain a performance evaluation result;
[0120] S1034. Update the parameters of the fine-tuned diffusion model according to the performance evaluation result to obtain an image generation model.
[0121] Specifically, the fine-tuned SD model is loaded for AI painting testing. After the model fine-tuning training is completed, the model weights will be saved in the set output_dir path. Next, using Stable Diffusion WebUI as a framework, the trained SD model is loaded for AI painting, and then manual evaluation and automatic evaluation are carried out. The evaluation metrics are mainly based on calculating the Fréchet Inception Distance (FID value) and CLIP score for the test set images. The parameters of the SD model are updated and adjusted through the evaluation feedback problems to obtain a trained image generation model.
[0122] Among them, FID (Fréchet inception distance) represents the similarity between the generated image and the real image, that is, the authenticity of the image. The closer the Fréchet Inception Distance is (the smaller the value), the better the effect of the generation model, that is, the image has high clarity and rich diversity. The CLIP score is used to evaluate the matching degree between the generated image and the input Prompt text and between the generated image and the input original image in text-to-image / image-to-image. The higher the CLIP score, the higher the matching degree between the generated image and the text description, that is, the image content is more relevant to the text description; the lower the CLIP score, the lower the matching degree between the generated image and the text description, that is, the image content is not relevant to the text description.
[0123] S104. Obtain user demand information, input the user demand information into the image generation model, and obtain a target image.
[0124] In some optional embodiments, the user demand information can be obtained in various ways. For example, the text description input by the user through the graphical user interface (GUI) and the selection of preset options. Or it is obtained through externally connected sensors, such as accessing a camera and a microphone, and collecting the user's gestures or voice commands through the camera and the microphone. The obtained user demand information is input into the trained image generation model for conditional generation, and a high-definition, high-resolution target image of more than 2k is output.
[0125] In summary, the processing flow of the image generation method based on adaptive image coding in the embodiments of the present invention is as Figure 9 shown:
[0126] The first step is to collect multi-size high-definition images and construct an image data set;
[0127] The second step is to perform data augmentation on the images in the image data set to obtain a first training data set and increase the diversity of data categories;
[0128] The third step is to construct an adaptive image coding module, and through the adaptive image coding module, perform image segmentation and size scaling on the first training data set, so that the model can have more sizes and higher-definition resolutions for training;
[0129] In the fourth step, input the second training dataset into the SD model for full-parameter fine-tuning training to achieve multi-size and high-resolution learning and improve the generalization ability of the model;
[0130] In the fifth step, use an independent test set to evaluate the generated image effects predicted by the SD model after fine-tuning training, and adjust the parameters of the SD model according to the evaluation feedback to ensure the generation of high-definition pictures and content authenticity, and obtain a trained image generation model;
[0131] In the sixth step, use the image generation model to generate images according to the user requirement information input by the user, and output a target image with high definition and high resolution.
[0132] The above describes the image generation method based on adaptive image coding in the embodiments of the present invention. It can be recognized that compared with the image generation models in the prior art, the embodiments of the present invention have the following advantages:
[0133] First, by adding a pre-adaptive image coding module during the model training process, the filtering of large-size images in the training dataset is reduced, and the image size scaling operation is also avoided, reducing the loss of global and local detail information of the pictures. Enable the image generation model to perform multi-size and high-resolution learning, improve the generalization ability of the model, and generate a target image with higher definition and higher resolution through the image generation model.
[0134] Second, through the full-parameter fine-tuning of the image generation model by combining the pre-trained SD model with the adaptive visual coding algorithm. Enable the model to perform continuous extended optimization learning and training of multi-size and high definition resolution. Enable the model itself to have the ability to generate high-definition resolution pictures, generate pictures with higher definition above 2k and more vivid content details, without relying on post-generation high-definition repair or other post-processing means to enhance the picture effects.
[0135] Referring to Figure 10 , the embodiments of the present invention also provide an image generation system based on adaptive image coding, including:
[0136] A data augmentation module for constructing an image dataset and performing data augmentation on the image dataset to obtain a first training dataset;
[0137] An adaptive processing module for constructing an adaptive image coding module and performing image segmentation on the first training dataset through the adaptive image coding module to obtain a second training dataset;
[0138] A model training module for training a preset diffusion model according to the second training dataset to obtain an image generation model;
[0139] An image generation model is used to obtain user requirement information, and the user requirement information is input into the image generation model to obtain a target image.
[0140] The content in the embodiments of the above image generation method based on adaptive image coding is applicable to the embodiments of the present image generation system based on adaptive image coding. The functions specifically implemented by the embodiments of the present image generation system based on adaptive image coding are the same as those of the embodiments of the above image generation method based on adaptive image coding, and the beneficial effects achieved are also the same as those of the embodiments of the above image generation method based on adaptive image coding.
[0141] Embodiments of the present invention also provide an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the above image generation method based on adaptive image coding. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0142] As Figure 11 shown is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present invention. Referring to Figure 11 , embodiments of the present invention provide an electronic device, including:
[0143] A processor 1001 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0144] A memory 1002 can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the image generation method based on adaptive image coding of the embodiments of the present invention;
[0145] An input / output interface 1003 is used to implement information input and output;
[0146] A communication interface 1004, which is used to implement the communication interaction between this device and other devices, can achieve communication through wired means (such as USB, network cable, etc.), or can also achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0147] A bus 1005, which transmits information between various components of the device (such as a processor 1001, a memory 1002, an input / output interface 1003, and a communication interface 1004);
[0148] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.
[0149] The embodiment of the present invention also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned image generation method based on adaptive image coding.
[0150] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] The embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0152] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0153] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0154] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a portable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0155] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0156] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0157] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0158] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0159] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0160] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. An image generation method based on adaptive image coding, characterized in that, It includes the following steps: Construct an image dataset, perform data augmentation on the image dataset to obtain a first training dataset; Construct an adaptive image coding module, and perform image segmentation on the first training dataset through the adaptive image coding module to obtain a second training dataset; Train a preset diffusion model according to the second training dataset to obtain an image generation model; Obtain user demand information, input the user demand information into the image generation model to obtain a target image.
2. The image generation method based on adaptive image coding according to claim 1, characterized in that The construction of the image dataset specifically includes: Collect multiple types of first training images, screen the first training images according to preset image screening conditions to obtain second training images; Filter the second training images to obtain third training images; Annotate the third training images to obtain fourth training images; Divide the fourth training images into a training set, a validation set, and a test set to obtain the image dataset.
3. The image generation method based on adaptive image coding according to claim 1, characterized in that, The step of performing image segmentation on the first training dataset through the adaptive image coding module to obtain a second training dataset specifically includes: Perform subgraph segmentation on the first training dataset to obtain subgraph position encoding information and subgraph size information; Perform size scaling on the first training dataset to obtain the original image size information; Concatenate the subgraph position encoding information, the subgraph size information, and the original image size information to obtain the second training dataset.
4. An image generation method based on adaptive image coding according to claim 3, characterized in that, The first training dataset includes multiple original images. The step of performing subgraph segmentation on the first training dataset to obtain subgraph position encoding information and subgraph size information specifically includes: Determine the target image size and the target aspect ratio, and calculate the original image size of the original image; Determine the initial segmentation times according to the original image size and the target image size, and then generate a candidate segmentation times set according to the initial segmentation times; Determine the segmentation rule, and obtain a list of candidate segmentation schemes according to the candidate segmentation times set and the segmentation rule. The list of candidate segmentation schemes includes multiple subgraph aspect ratios; Calculate the proximity ratio of each subgraph aspect ratio to the target aspect ratio, and then determine the best subgraph aspect ratio according to the proximity ratio; Segment each original image according to the best subgraph aspect ratio and the original image size to obtain the subgraph position encoding information and the subgraph size information.
5. A method for generating an image based on adaptive image coding according to claim 4, wherein The step of segmenting each original image according to the best subgraph aspect ratio and the original image size to obtain the subgraph position encoding information and the subgraph size information specifically includes: Calculate the subgraph size information according to the original image size and the best subgraph aspect ratio; Segment the original image according to the subgraph size information to obtain multiple subgraphs; Calculate the position coordinates of each subgraph in the original image to obtain the subgraph position encoding information.
6. The image generation method based on adaptive image coding according to claim 3, wherein, The step of concatenating the subgraph position encoding information, the subgraph size information, and the original image size information to obtain the second training dataset specifically includes: Convert the sub - graph position encoding information, the sub - graph size information, and the original image size information into Fourier encoding; Concatenate each of the Fourier encodings to obtain the second training dataset.
7. A method for generating an image based on adaptive image coding according to claim 1, wherein The diffusion model includes a U - Net module. Training a preset diffusion model according to the second training dataset to obtain an image generation model specifically includes: Set training parameters and freeze the weights of the U - Net module; Input the second training dataset into the diffusion model and perform fine - tuning training on the diffusion model according to the training parameters; Evaluate the performance of the fine - tuned diffusion model according to the Fréchet inception distance and the CLIP score to obtain a performance evaluation result; Update the parameters of the fine - tuned diffusion model according to the performance evaluation result to obtain the image generation model.
8. An image generation system based on adaptive image coding, characterized in that, Including: A data augmentation module for constructing an image dataset and performing data augmentation on the image dataset to obtain a first training dataset; An adaptive processing module for constructing an adaptive image encoding module and performing image segmentation on the first training dataset through the adaptive image encoding module to obtain a second training dataset; A model training module for training a preset diffusion model according to the second training dataset to obtain the image generation model; An image generation module for obtaining user requirement information and inputting the user requirement information into the image generation model to obtain a target image.
9. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it implements the steps of the image generation method based on adaptive image encoding according to any one of claims 1 to 7.
10. A storage medium, the storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the image generation method based on adaptive image encoding according to any one of claims 1 to 7.