An image generation method, controller and storage medium based on a diffusion model
Through the image generation method based on the diffusion model, the urban scene image data set is classified and trained, which solves the shortcomings of Dreambooth in the learning of complex urban scene features, improves the diversity and accuracy of image features, and the generated images show good expression effects.
Patent Information
- Application Number
- CN202510199745.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Dreambooth performs poorly in feature learning with complex elements and diverse goals in urban scenes, and lacks generalization ability, resulting in the lack of diversity in generated images, overfitting, low accuracy in feature learning, and poor expression effect.
Using an image generation method based on diffusion model, by classifying and processing the urban scene image data set, different batches of images are obtained, and iteratively trained with different learning rates and learning frequency to optimize model parameters to better capture image features.
The feature diversity and feature learning accuracy of urban scene images are improved, so that the generated feature images have excellent expression effects, and the model convergence level and generation accuracy performance are improved.
Smart Images

Figure CN119693487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an image generation method, a controller, and a storage medium based on a diffusion model. Background Art
[0002] Urban scenes contain various elements, such as buildings, vehicles, pedestrians, and animals. The elements of urban scene images are diverse and complex, making the feature learning of urban scene images a difficult problem. Dreambooth is a personalized image generation technology based on deep learning. Dreambooth realizes the learning of a certain feature object through small-sample training. Such a training method is easy to learn some object features, so as to generate high-quality images containing target features, enabling Dreambooth to perform well in the single-building and small-sample graphic training. However, for group buildings, Dreambooth is not good at feature learning with complex elements and diverse targets like urban scenes. Its generalization ability may be insufficient, resulting in a lack of diversity in the generated images, and there will also be a potential phenomenon of overfitting to specific objects, making the generated images single, with unclear primary and secondary, resulting in low accuracy of feature learning of urban scene images and poor expression effects of urban scene images. Summary of the Invention
[0003] This application aims to solve at least one of the technical problems existing in the prior art. To this end, embodiments of this application provide an image generation method, a controller, and a storage medium based on a diffusion model, which are beneficial to improving the feature diversity of urban scene images and the accuracy of scene image feature learning, and enabling scene images to have excellent expression effects.
[0004] In a first aspect, embodiments of this application provide an image generation method based on a diffusion model, including:
[0005] Obtain an urban scene image dataset, and perform classification processing on the urban scene image dataset to obtain a first batch of urban scene images, a second batch of urban scene images, and a third batch of urban scene images;
[0006] Input the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images into a diffusion model for iterative training processing to obtain a trained diffusion model;
[0007] Obtain an urban scene image to be processed, and input the urban scene image to be processed into the trained diffusion model to obtain a feature image corresponding to the urban scene image to be processed.
[0008] According to some embodiments of the present application, inputting the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images into the diffusion model for iterative training processing includes:
[0009] Performing paired training on the first batch of urban scene images at a first learning rate and a first learning frequency, so that the first batch of urban scene images form a mapping relationship between prompt words and various scene images;
[0010] Performing feature sampling training on the second batch of urban scene images at a second learning rate and a second learning frequency, so that the features of the second batch of urban scene images are diversified;
[0011] Performing reinforcement training on the third batch of urban scene images at a third learning rate and a third learning frequency, so that the third batch of urban scene images form a one-to-one mapping relationship between feature words and corresponding types of scene images.
[0012] According to some embodiments of the present application, the first batch of urban scene images are multi-viewpoint urban scene images without screening, the second batch of urban scene images are high-definition urban scene images with a resolution higher than a first preset resolution, and the third batch of urban scene images are scene images of a preset theme special atlas.
[0013] According to some embodiments of the present application, after obtaining the urban scene image dataset, the method further includes:
[0014] Performing text annotation processing on the urban scene image dataset to form a mapping relationship between the text annotation and the urban scene image dataset;
[0015] Uploading the mapping relationship between the text annotation and the urban scene image dataset to a terminal device, so that an auditor can judge the accuracy of the mapping relationship between the text annotation and the urban scene image dataset.
[0016] According to some embodiments of the present application, the performing reinforcement training on the third batch of urban scene images at a third learning rate and a third learning frequency includes:
[0017] Performing reinforcement training on the third batch of urban scene images based on the diffusion model at a third learning rate and a third learning frequency;
[0018] Binding a certain specific type of scene image to one or several feature words through a regularization training method, and generating corresponding keywords;
[0019] When responding to the keyword, generating a scene image corresponding to the type of the keyword.
[0020] According to some embodiments of the present application, determine the path of the third batch of urban scene images, and regularize the path of the third batch of urban scene images and the regularization prior loss weight;
[0021] Set the resolution of the third batch of urban scene images, and enable the arb bucket in Stable Diffusion to allow the third batch of urban scene images with non-fixed aspect ratios;
[0022] Set the minimum resolution of the arb bucket to the first threshold, the maximum resolution to the second threshold, and the resolution division unit of the arb bucket to the third threshold;
[0023] Set the number of training epochs of the diffusion model to the fourth threshold, and set the training batch size to the fifth threshold.
[0024] According to some embodiments of the present application, the first learning rate is less than the second learning rate, the second learning rate is lower than the third learning rate, the first learning frequency is lower than the second learning frequency, and the second learning frequency is lower than the third learning frequency.
[0025] According to some embodiments of the present application, use a first optimizer to optimize the first batch of urban scene images and the second batch of urban scene images;
[0026] Use a second optimizer to optimize the third batch of urban scene images.
[0027] In a second aspect, an embodiment of the present application provides a controller, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor runs the computer program, it executes the method described in the technical solution of the first aspect above.
[0028] In a third aspect, a computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method described in the technical solution of the first aspect above.
[0029] One of the advantages or beneficial effects that the image generation method, controller, and storage medium based on the diffusion model provided by the embodiments of the present application at least have is as follows: Collect an image dataset containing different urban scenes, improve the diversity of the urban scene image dataset, classify the obtained urban scene image dataset according to content, style, and features, and obtain the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images in different batches, and use them as the training dataset of the diffusion model, and input them into the diffusion model for iterative training. By learning the distribution of the urban scene image data in different batches, the diffusion model gradually optimizes the model parameters to better capture the features of the urban scene images, making the features of the urban scene images richer and more distinct between primary and secondary, thereby improving the feature diversity of the urban scene images and the accuracy of scene image feature learning, obtaining a trained diffusion model, improving the convergence level of the trained diffusion model, and greatly improving the accuracy performance of scene image generation. When the urban scene image to be processed is input into the trained diffusion model for further processing or analysis, the trained diffusion model can quickly learn the features of the input urban scene image to be processed and generate a corresponding feature image, and the generated feature image has an excellent expression effect.
[0030] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings
[0031] Figure 1 is a flowchart of an image generation method based on a diffusion model provided by an embodiment of the present application;
[0032] Figure 2 is a flowchart of a method for iterative training of a diffusion model provided by an embodiment of the present application;
[0033] Figure 3 is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application;
[0034] Figure 4 is a flowchart of a method for intensively training the third batch of urban scene images at a third learning rate and a third learning frequency provided by an embodiment of the present application;
[0035] Figure 5 is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application;
[0036] Figure 6It is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application;
[0037] Figure 7 It is a schematic structural diagram of a controller provided by an embodiment of the present application. Detailed implementation manners
[0038] This part will describe in detail the specific embodiments of the present application. The preferred embodiments of the present application are shown in the drawings. The function of the drawings is to supplement the description in the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present application, but it cannot be understood as a limitation on the protection scope of the present application.
[0039] In the description of the present application, the meaning of "several" is one or more, the meaning of "multiple" is more than two, understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, understandings such as "above", "below", "within", etc. include the recited number, "any one" means one or more, and expressions such as "at least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. If there is a description of first and second, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0040] It should be noted that words such as "set", "installed", "connected", etc. in the embodiments of the present application should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the embodiments of the present application in combination with the specific content of the technical solution. For example, the term "connected" can be a mechanical connection, an electrical connection, or a connection that can communicate with each other; it can be directly connected or indirectly connected through an intermediate medium.
[0041] It should be noted that the technical features involved in the various implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.
[0042] Urban scenes contain various elements, such as buildings, vehicles, pedestrians, and animals, etc. The elements of urban scene images have the characteristics of diversity and complexity, making the feature learning of urban scene images a difficult problem.
[0043] Dreambooth is a deep learning-based personalized image generation technology. Dreambooth learns a certain characteristic object through small sample training. Such a training method is easy to learn some object features, thereby generating high-quality images containing target features, enabling Dreambooth to perform well in the single-building small sample graphic training. However, for group buildings, Dreambooth is not good at learning features with complex elements and diverse targets like urban scenes. Its generalization ability may be insufficient, resulting in a lack of diversity in the generated images and potentially overfitting to specific objects, making the generated images single, with unclear primary and secondary elements, leading to low accuracy in feature learning of urban scene images and poor expression effects of urban scene images.
[0044] Based on this, the embodiments of the present application provide an image generation method, a controller, and a storage medium based on a diffusion model, which are beneficial to improving the feature diversity of urban scene images and the accuracy of scene image feature learning, enabling scene images to have excellent expression effects.
[0045] Next, please refer to the attached drawings to further elaborate on the image generation method, controller, and storage medium based on the diffusion model provided by the embodiments of the present application.
[0046] Refer to Figure 1 as shown in Figure 1 is a flowchart of an image generation method based on a diffusion model provided by the embodiments of the present application. The image generation method based on the diffusion model includes but is not limited to steps S100 to S300. Specifically,
[0047] Step S100: Obtain an urban scene image dataset, and perform classification processing on the urban scene image dataset to obtain the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images;
[0048] Step S200: Input the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images into the diffusion model for iterative training processing to obtain a trained diffusion model;
[0049] Step S300: Obtain an urban scene image to be processed, input the urban scene image to be processed into the trained diffusion model, and obtain a feature image corresponding to the urban scene image to be processed.
[0050] In some embodiments of the present application, the image generation method based on a diffusion model includes: collecting an image dataset containing different urban scenes. The image dataset of urban scenes may be from public datasets, satellite images, or street view cameras to improve the diversity of the image dataset of urban scenes. Classifying the obtained image dataset of urban scenes, categorizing it according to the content, style, and features of the urban scene images to obtain the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images. The different batches of urban scene images are used as the training dataset for the diffusion model. Inputting the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images into the diffusion model for iterative training. The diffusion model learns the distribution of the image data of different batches of urban scene images and gradually optimizes the model parameters to better capture the features of the urban scene images, making the features of the urban scene images richer and more distinct between the primary and secondary, thereby improving the feature diversity of the urban scene images and the accuracy of learning the scene image features, obtaining a trained diffusion model. The convergence level of the trained diffusion model is improved, and the accuracy performance of scene image generation is also greatly improved. Inputting the urban scene image to be processed into the trained diffusion model for further processing or analysis of the urban scene image. The trained diffusion model can quickly learn the features of the input urban scene image to be processed and generate a corresponding feature image, and the generated feature image has an excellent expression effect.
[0051] In some embodiments of the present application, the first batch of urban scene images are multi-viewpoint urban scene images without screening, the second batch of urban scene images are high-definition urban scene images with a resolution higher than the first preset resolution, and the third batch of urban scene images are scene images of a preset theme special atlas.
[0052] Urban scenes contain various elements such as buildings, vehicles, pedestrians, and animals. The elements of urban scene images have the characteristics of diversity and complexity. In the embodiments of the present application, the obtained image dataset of urban scenes is classified, and a large number of image datasets of urban scenes are disassembled into multi-type sub-datasets according to the standards and requirements of content, style, and features to obtain the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images. By classifying the image dataset of urban scenes, it can be ensured that the image dataset of urban scenes serves various application requirements more orderly and efficiently, and at the same time, it is also convenient for subsequent data analysis and diffusion model training. In addition, the processed scene image dataset is effectively stored and managed to ensure the accessibility and security of the data. And during the processing, quality control needs to be carried out on the scene image data to ensure the accuracy and usability of the images.
[0053] In some embodiments, the first batch of urban scene images are unfiltered multi-view urban scene images, which contain urban landscapes with various perspectives and resolutions. The second batch of urban scene images are high-definition urban scene images with a resolution higher than the first preset resolution. The second batch of urban scene images have higher clarity and details and are suitable for application scenarios that require high-resolution images. The third batch of urban scene images are scene images of a preset theme special atlas, including but not limited to bird's-eye view theme special atlas, semi-bird's-eye view theme special atlas, building monomer theme special atlas, park theme special atlas, waterfront scene theme special atlas, city center theme special atlas, etc. Through the third batch of urban scene images, the special atlas can be screened according to specific subjects or uses.
[0054] By classifying and processing the urban scene image dataset to obtain different batches of urban scene images, it can ensure that the urban scene image dataset is more orderly, facilitate better capturing the features of urban scene images, make the features of urban scene images more abundant, with distinct primary and secondary features, improve the feature diversity of urban scene images and the accuracy of scene image feature learning, and at the same time facilitate subsequent data analysis and diffusion model training.
[0055] Refer to Figure 2 shown in Figure 2 is a flowchart of a method for iterative training processing of a diffusion model provided by an embodiment of the present application. The method for iterative training processing of the diffusion model includes but is not limited to steps S210 to S230. Specifically,
[0056] Step S210: Pairwise train the first batch of urban scene images with the first learning rate and the first learning frequency so that the first batch of urban scene images form a mapping relationship between prompt words and various scene images;
[0057] Step S220: Perform feature sampling training on the second batch of urban scene images with the second learning rate and the second learning frequency so that the features of the second batch of urban scene images are diversified;
[0058] Step S230: Perform reinforcement training on the third batch of urban scene images with the third learning rate and the third learning frequency so that the third batch of urban scene images form a mapping relationship in which feature words correspond one-to-one with corresponding types of scene images.
[0059] In some embodiments of the present application, after obtaining an urban scene image dataset and classifying the urban scene image dataset to obtain the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images, the first batch of urban scene images, the second batch of urban scene images, and the third batch of urban scene images are input into a diffusion model for iterative training to obtain a diffusion model with a higher convergence level and higher accuracy in generating scene images.
[0060] The method for iterative training of the diffusion model includes: rapidly learning the first batch of urban scene images at a first learning rate and a first learning frequency, enabling the prompt words to form a mapping relationship with various scene images, which helps the diffusion model understand the content and context of the scene images, so as to better utilize this information in feature image generation or classification tasks; it can be understood that the number of repeated trainings for each urban scene image in the training of the first batch of urban scene images is relatively small, generally 5 - 10 times, and the first learning rate is also relatively low, trying to cover the rapid learning of all urban scene images and allowing the prompt words to form a fuzzy mapping relationship with various images. Sampling and training the features of the second batch of urban scene images at a second learning rate and a second learning frequency to diversify the features of the second batch of urban scene images. This process may include using a convolutional neural network or other feature extraction techniques to identify and learn the key features in the urban scene images. The goal is to diversify the features of the second batch of urban scene images, which helps improve the recognition and understanding ability of the diffusion model for different types of urban scene images; it can be understood that the second batch of urban scene images mainly trains high-resolution urban scene images. The second learning rate for each urban scene image is relatively high but the number of repeated learning times is relatively small (i.e., the second learning frequency is small), preventing overfitting caused by a large number of local sampling repeated learnings of high-resolution urban scene images and enabling the diffusion model to learn a certain method for depicting details; strengthening the training of the third batch of urban scene images at a third learning rate and a third learning frequency, and through reinforcement learning, ensuring that the diffusion model can form a one-to-one mapping relationship between the feature words of the third batch of urban scene images and the corresponding type of scene images. This mapping relationship is crucial for the performance of the diffusion model in specific tasks such as image retrieval or scene classification; it can be understood that the third batch of urban scene images mainly strengthens the learning of special word semantics, binding a specific type of urban scene image to a certain word or several words through a regularization training method, and can be activated by specific keywords when needed.
[0061] The urban scene images of different batches are trained through a bucketing mode, enabling the model to gradually learn and adapt to complex tasks. Meanwhile, ensuring that each stage has clear goals and optimization directions, gradually optimizing the model parameters to better capture the features of urban scene images, making the features of urban scene images richer, with distinct primary and secondary features, thereby improving the feature diversity of urban scene images and the accuracy of scene image feature learning, obtaining a trained diffusion model, improving the convergence level of the trained diffusion model, and greatly improving the performance of scene image generation accuracy.
[0062] In some embodiments of the present application, the first learning rate is less than the second learning rate, the second learning rate is lower than the third learning rate, the first learning frequency is lower than the second learning frequency, and the second learning frequency is lower than the third learning frequency.
[0063] The method for iterative training of the diffusion model includes paired training of the first batch of urban scene images at the first learning rate and the first learning frequency to form a mapping relationship between the prompt words and various scene images; feature sampling training of the second batch of urban scene images at the second learning rate and the second learning frequency to diversify the features of the second batch of urban scene images; and reinforcement training of the third batch of urban scene images at the third learning rate and the third learning frequency to form a one-to-one mapping relationship between the feature words and the corresponding types of scene images. Among them, the first learning rate is less than the second learning rate, the second learning rate is lower than the third learning rate, the first learning frequency is lower than the second learning frequency, and the second learning frequency is lower than the third learning frequency.
[0064] Training the first batch of urban scene images with a lower first learning rate and a lower first learning frequency can achieve fast learning that covers all urban scene images as much as possible, and also helps the diffusion model to learn stably, avoiding the divergence of the model caused by an overly large update step size. As the training progresses, gradually increasing the learning rate can accelerate the convergence speed of the model, which helps the diffusion model to learn better parameter settings faster. Therefore, using a second learning rate slightly higher than the first learning rate and a learning frequency slightly higher than the first learning rate to perform feature sampling training on the second batch of urban scene images can better capture the features of urban scene images, making the features of urban scene images more abundant, with distinct primary and secondary features, thereby improving the feature diversity of urban scene images and the accuracy of scene image feature learning. Reducing the learning rate in the later stage of training can reduce the risk of overfitting of the diffusion model on the urban scene image dataset and improve the generalization ability of the diffusion model. Therefore, using a relatively high third learning rate and a third learning frequency to perform intensive training on the third batch of urban scene images, with the number of various preset theme special picture sets ranging from 50 to 100 images for intensive training, can ensure through intensive training that the diffusion model can form a one-to-one mapping relationship between the feature words of the third batch of urban scene images and the corresponding type of scene images, improving the accuracy of the generated feature images.
[0065] It should be noted that in some embodiments of the present application, a teacher-student framework is used to train different batches of urban scene images. In deep learning, learning_rate (learning rate) is a key hyperparameter that controls the weight update step size of the diffusion model, and it determines the magnitude of the diffusion model parameter update in each iteration. learning_rate_te (teacher learning rate) is usually used in the teacher-student learning framework, where the parameters of the teacher model are updated at a slower speed to stabilize the learning process of the student.
[0066] Training the first batch of urban scene images with the first learning rate and the first learning frequency includes:
[0067] Set learning_rate = 0.000001, which is a very small learning rate, equivalent to making only a tiny adjustment to the weights in each iteration. This small learning rate helps the diffusion model to converge stably during training, avoiding instability or divergence in the training process caused by an overly large step size.
[0068] Set learning_rate_te = 0.0000005. This learning rate is used for updating the parameters of the teacher model. Since the value of learning_rate_te is smaller than the learning rate of learning_rate, it allows the parameters of the teacher model to be updated more slowly, thereby providing smoother gradient information for the student model.
[0069] Perform feature sampling training on the second batch of urban scene images with the second learning rate and the second learning frequency, including:
[0070] Set learning_rate = 0.000005. This value is higher than the first learning rate. Using the second learning rate higher than the first learning rate helps the model converge stably during training. Especially in the initial stage of training, it can avoid instability or divergence caused by too large a step size;
[0071] Set learning_rate_te = 0.000001. This learning_rate_te is smaller than the learning_rate, making the parameter update speed of the teacher model slower than that of the student model. This setting can help stabilize the learning process of the student because the teacher model provides a smoother gradient information, which helps the student model converge faster while maintaining the stability of the training process.
[0072] Perform reinforcement training on the third batch of urban scene images with the third learning rate and the third learning frequency, including:
[0073] Set learning_rate = 0.000005. This is a relatively high learning rate that can strengthen the learning of special word semantics. By means of regularization training, a specific type of image is bound to a certain word or several words, so that a one-to-one mapping relationship between the feature words and the corresponding type of scene images is formed for the third batch of urban scene images;
[0074] Set learning_rate_te = 0.000001, which can finely adjust the weights in the later stage of the training process when the model is approaching convergence, improving the model performance.
[0075] Refer to Figure 3 as shown Figure 3 is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application. The image generation method based on the diffusion model includes but is not limited to steps S110 to S120. Specifically,
[0076] Step S110: Perform text annotation processing on the urban scene image dataset to form a mapping relationship between the text annotation and the urban scene image dataset;
[0077] Step S120: Upload the mapping relationship between the text annotation and the urban scene image dataset to the terminal device so that the reviewer can judge the accuracy of the mapping relationship between the text annotation and the urban scene image dataset.
[0078] In some embodiments of the present application, the image generation method based on the diffusion model further includes: after obtaining the urban scene image dataset, preprocessing the urban scene image dataset. By annotating the urban scene image dataset with text, that is, using text to explain the content of the urban scene image, a mapping relationship between the text annotation and the urban scene image dataset is formed, and a structured urban scene image dataset is formed, which is convenient for the diffusion model to recognize and understand the content of the urban scene image. The mapping relationship between the processed text annotation and the urban scene image dataset is uploaded to the terminal device or the server for easy management and distribution. The reviewer accesses the uploaded mapping relationship through the terminal device to check whether the mapping relationship between the text annotation and the urban scene image dataset is consistent, complete, and accurate. When the reviewer reviews that the mapping relationship between the text annotation and the urban scene image dataset is accurate, the mapping relationship between the text annotation and the urban scene image dataset is stored for subsequent model training and analysis; when the reviewer reviews that the mapping relationship between the text annotation and the urban scene image dataset is inaccurate, the text annotation is adjusted or modified to make the text annotation correspond to the urban scene image dataset, ensuring the accuracy of the mapping relationship between the text annotation and the urban scene image dataset and improving the quality of the urban scene image dataset.
[0079] By performing annotation processing on the urban scene image dataset after obtaining it, accurate training data can be provided for subsequent machine learning and deep learning training, which is beneficial to accurately identifying and learning the characteristics of the urban scene.
[0080] Refer to Figure 4 as shown Figure 4 is a flowchart of a method for intensively training the third batch of urban scene images at a third learning rate and a third learning frequency provided by an embodiment of the present application. The method for intensively training the third batch of urban scene images at a third learning rate and a third learning frequency includes but is not limited to steps S231 to S232. Specifically,
[0081] Step S231: Intensively train the third batch of urban scene images based on the diffusion model at a third learning rate and a third learning frequency;
[0082] Step S232: Bind a specific type of scene image to one or several feature words through a regularization training method and generate corresponding keywords;
[0083] Step S233: When responding to the keyword, generate a scene image of the type corresponding to the keyword.
[0084] In some embodiments of the present application, the method for enhancing the training of the third batch of urban scene images with a third learning rate and a third learning frequency includes: enhancing the training of the third batch of urban scene images based on a diffusion model with the third learning rate and the third learning frequency. It should be noted that the third learning rate is a relatively high learning rate, and the third learning frequency is a relatively high learning frequency. By setting a high learning rate and a high learning frequency to enhance the training of the third batch of urban scene images, special learning can be better carried out. Through a regularization training method, a specific type of scene image is bound to a certain feature word or several feature words to form a mapping relationship between the feature word and the specific type of scene image, and corresponding keywords are generated according to the mapping relationship between the feature word and the specific type of scene image. When responding to the keyword, a scene image of the corresponding type to the keyword is generated, thereby obtaining a feature image corresponding to the keyword, improving the accuracy of scene image feature learning, and enabling the scene image to have an excellent expression effect.
[0085] Referring Figure 5 as shown Figure 5 is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application. The image generation method based on the diffusion model includes but is not limited to steps S234 to S237. Specifically,
[0086] Step S234: Determine the path of the third batch of urban scene images, and regularize the path of the third batch of urban scene images and the regularization prior loss weight;
[0087] Step S235: Set the resolution of the third batch of urban scene images, and enable the arb bucket in Stable Diffusion to allow the third batch of urban scene images with non-fixed aspect ratios;
[0088] Step S236: Set the minimum resolution of the arb bucket as the first threshold, the maximum resolution as the second threshold, and set the resolution division unit of the arb bucket as the third threshold;
[0089] Step S237: Set the number of training rounds of the diffusion model as the fourth threshold, and set the training batch size as the fifth threshold.
[0090] In some embodiments of the present application, before enhancing the training of the third batch of urban scene images with the third learning rate and the third learning frequency, it also includes parameter setting for the diffusion model so that the parameters of the diffusion model are suitable for training the third batch of urban scene images. The method for setting the parameters of the diffusion model includes: First, determine the path of the third batch of urban scene images, determine a suitable folder to store the third batch of urban scene images, and determine the storage or access path for the third batch of urban scene images so that the diffusion model can correctly load and process the third batch of urban scene images.
[0091] Apply regularization techniques to adjust the distribution of the urban scene image data in the third batch to have zero mean and unit variance, which helps the stability and efficiency of diffusion model training. Set the prior loss weight to balance the distribution of different classes in the training data and prevent the diffusion model from biasing towards the majority class. Set an appropriate resolution for the urban scene images in the third batch. By setting the [width, height] resolution of the urban scene images in the third batch to [512, 512], the quality of the generated feature images and the processing ability of the diffusion model can be improved. Enable the arb bucket in Stable Diffusion to allow the urban scene images in the third batch with non-fixed aspect ratios, support images with non-fixed aspect ratios, and improve the flexibility and applicability of the diffusion model. When setting the resolution of the urban scene images in the third batch, non-square resolutions are supported, but the resolution must be a multiple of 64. Set the minimum resolution of the arb bucket as the first threshold, the maximum resolution as the second threshold, and the resolution division unit of the arb bucket as the third threshold. In one embodiment, the first threshold is 256, the second threshold is 1024, and the third threshold is 64. By setting the minimum resolution of the arb bucket to 256, the maximum resolution to 1024, and the resolution division unit of the arb bucket to 64 to process the urban scene images in the third batch, the accuracy of scene image feature learning can be improved, making the scene images have excellent expression effects. Set the number of training rounds of the diffusion model as the fourth threshold to determine the total number of training rounds of the diffusion model, thereby controlling the training cycle and performance level of the diffusion model. Set the training batch size as the fifth threshold to determine the number of images processed in each training iteration and the use of memory and computing resources. By setting the fourth threshold to 10 and the fifth threshold to 1, that is, setting the number of training rounds of the diffusion model to 10 and the training batch size to 1 for feature learning, the generated feature images can be more rich, with distinct primary and secondary, and have better expression effects.
[0092] By setting the parameters of the diffusion model, it can be ensured that the urban scene images in the third batch are effectively processed and trained in the diffusion model. This training process allows the diffusion model to learn the specific features of the urban scene and generate high-quality feature image outputs.
[0093] It should be noted that in practical applications, the settings of these parameters such as the first threshold, the second threshold, the third threshold, the fourth threshold, and the fifth threshold need to be adjusted according to the specific tasks and the characteristics of the urban scene image dataset to achieve the best training effect. The embodiments of the present application do not limit the sizes of the first threshold, the second threshold, the third threshold, the fourth threshold, and the fifth threshold.
[0094] Refer to Figure 6 as shown Figure 6It is a flowchart of an image generation method based on a diffusion model provided by another embodiment of the present application. The image generation method based on the diffusion model includes but is not limited to steps S400 to S410. Specifically,
[0095] Step S400: Optimize the first batch of urban scene images and the second batch of urban scene images using a first optimizer;
[0096] Step S410: Optimize the third batch of urban scene images using a second optimizer.
[0097] In some embodiments of the present application, during the training of different batches of urban scene images based on the diffusion model, the first optimizer is used to optimize the first batch of urban scene images and the second batch of urban scene images. The first optimizer is the AdamW8bit optimizer.
[0098] The AdamW8bit optimizer is a low-precision optimizer that can reduce memory consumption and computational costs. It reduces memory occupancy by reducing the bit width of weights and gradients, thus allowing for the training of larger diffusion models with limited GPU memory. In addition, the AdamW8bit optimizer can improve stability through automatic adaptive clipping and block-wise quantization. By optimizing the first batch of urban scene images and the second batch of urban scene images using the AdamW8bit optimizer, the efficiency and performance of model training are improved through the low memory occupancy and fast optimization characteristics of the AdamW8bit optimizer.
[0099] The third batch of urban scene images is optimized using a second optimizer. The second optimizer is the SGDesterov8bit optimizer. The SGDesterov8bit optimizer provides the same interface as the standard 32-bit optimizer but has less memory occupancy and faster optimization speed. The SGDesterov8bit is particularly suitable for training or fine-tuning large models on memory-constrained GPUs. By optimizing the third batch of urban scene images using the SGDesterov8bit optimizer, the efficiency and performance of model training can be improved, a diffusion model with a higher convergence level and higher accuracy performance in generating scene images can be obtained, ensuring the accuracy and availability of the urban scene image dataset, and providing support for subsequent analysis and applications.
[0100] Refer to Figure 7 , Figure 7It is a schematic structural diagram of a controller 1000 provided by an embodiment of the present application. It includes a processor 1001, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the methods provided by the embodiments of the present application; a memory 1002, which can be implemented in forms such as a read-only memory 1002 (ROM), a static storage device, a dynamic storage device, or a random access memory 1002 (Random Access Memory, RAM). The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called and executed by the processor 1001 for the embodiments of the present application; an input / output interface 1003, which is used to implement information input and output; a communication interface 1004, which is used to implement communication interaction between this device and other devices, and can achieve communication through a wired method (such as USB, network cable, etc.) or can also achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); a bus, which transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004); among which the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus.
[0101] Those of ordinary skill in the art will appreciate that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer-readable storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0102] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application can be realized and attained by the structure particularly pointed out in the specification, claims, and drawings.
Claims
1. An image generation method based on a diffusion model, characterized in that: include: Acquire a data set of urban scene images, and classify the data set of urban scene images to obtain a first batch of urban scene images, a second batch of urban scene images, and a third batch of urban scene images, wherein the first batch of urban scene images are unscreened multi-viewpoint urban scene images, the second batch of urban scene images are high-definition urban scene images having a resolution higher than a first preset resolution, and the third batch of urban scene images are scene images of a preset theme special atlas; The first batch of urban scene images, the second batch of urban scene images and the third batch of urban scene images are input into the diffusion model for iterative training processing to obtain a trained diffusion model, including: pairing training the first batch of urban scene images with a first learning rate and a first learning frequency, so that the first batch of urban scene images form a mapping relationship between prompt words and various scene images; feature sampling training is performed on the second batch of urban scene images with a second learning rate and a second learning frequency, so that the features of the second batch of urban scene images are diversified; reinforcement training is performed on the third batch of urban scene images with a third learning rate and a third learning frequency, so that the third batch of urban scene images form a one-to-one mapping relationship between feature words and corresponding types of scene images, specifically, reinforcement training is performed on the third batch of urban scene images based on the diffusion model with a third learning rate and a third learning frequency; a certain type of scene image is bound to one or several feature words by a regularized training method, and a corresponding keyword is generated; when responding to the keyword, a scene image of the type corresponding to the keyword is generated; Acquire a city scene image to be processed, input the city scene image to be processed into the trained diffusion model, and obtain a feature image corresponding to the city scene image to be processed. Specifically, determine the path of the third batch of city scene images, and regularize the path and regularization prior loss weight of the third batch of city scene images; set the resolution of the third batch of city scene images, and enable the arb bucket in Stable Diffusion to allow the third batch of city scene images with non-fixed aspect ratio; set the minimum resolution of the arb bucket to a first threshold, the maximum resolution to a second threshold, and the resolution division unit of the arb bucket to a third threshold; set the number of training rounds of the diffusion model to a fourth threshold, and set the training batch size to a fifth threshold.
2. The image generation method based on the diffusion model according to claim 1, characterized in that: After acquiring the urban scene image dataset, the method further includes: Performing text annotation processing on the urban scene image dataset to form a mapping relationship between the text annotation and the urban scene image dataset; The mapping relationship between the text annotation and the urban scene image data set is uploaded to a terminal device so that a reviewer can judge the accuracy of the mapping relationship between the text annotation and the urban scene image data set.
3. The image generation method based on the diffusion model according to claim 1, characterized in that: The first learning rate is smaller than the second learning rate, the second learning rate is lower than the third learning rate, the first learning frequency is lower than the second learning frequency, and the second learning frequency is lower than the third learning frequency.
4. The image generation method based on the diffusion model according to claim 1, characterized in that: Also includes: Using a first optimizer to optimize the first batch of urban scene images and the second batch of urban scene images; The third batch of urban scene images is optimized using a second optimizer.
5. A controller, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method according to any one of claims 1 to 4 when executing the computer program.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Model training method, building effect picture generation method, equipment and medium
CN117351325A
Medical spectrum reconstruction method and device
CN119027342A