Method for generating high-resolution water meter image based on text of diffusion model
By combining VAE autoencoder, CLIP model, diffusion denoising U-Net network, and combining ESRGAN network for super-resolution processing, the problems of blurred and detailed loss of water meter image generation in the prior art are solved, and high-resolution, clear and arbitrary resized image generation is achieved.
Patent Information
- Application Number
- CN202510205237.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems of blurring and loss of details when generating high-resolution water meter images, and the generated images are fixed in size, making it difficult to adjust at will, affecting the network operation speed.
A high-resolution image method based on diffusion model is adopted, combined with VAE autoencoder, CLIP model and diffusion denoising U-Net network, features are captured through multi-scale attention aggregation module to improve image generation quality. Super-resolution processing is used using the ESRGAN network to generate clear and resizable images.
It realizes the generation of high-resolution, clear and clear water meter images, and can adjust the image size without losing details, improving the network's running speed and diversity of image generation.
Smart Images

Figure CN119991444A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent generation technology of text-generated images, and in particular to a high-precision deep generative network model of a latent space suitable for multimodal scenarios, specifically a method for generating high-resolution images from text based on a diffusion model, and belongs to the field of computer vision and artificial intelligence generation. Background Art
[0002] So far, the review of water meter readings is an indispensable step, and relevant departments need to review water meter readings to calculate water consumption. At this stage, the traditional manual meter reading method is still the main method, which results in a large amount of manpower and material resources. In reality, many water meters are installed in very hidden locations, which makes the review difficult. With the maturity and development of artificial intelligence technology, automatic detection and recognition technology based on deep learning network models makes the intelligent needs of various fields possible. However, the premise of intelligence still requires some real water meters to produce data sets, which are sent to the deep network model for training, so that the network can read accurately. However, in reality, the types of water meters are complex and diverse, and some are extremely difficult to shoot, which makes the network have incomplete data sets during training, resulting in inaccurate intelligent detection of some water meters.
[0003] In order to increase the diversity of water meter datasets, text-generated image technology is a good method. The early development of text-generated image technology mainly focused on GAN (Generative Adversarial Network). Through adversarial training, a generative model and a discriminative model compete with each other to improve the quality and authenticity of the generated content. But the training process of GAN is a dynamic game process, where the generator and the discriminator compete with each other to improve their respective performance. However, this adversarial training method may lead to instability during the training process, such as mode collapse, gradient disappearance or explosion. These problems make GAN training difficult, and after real experiments, the generated water meter pictures are blurred, details are lost, and they do not have the diversified effect of water meter datasets.
[0004] In this case, a network model for text-generated images is urgently needed to generate various types of water meters. In the latest diffusion model, very good results have been achieved in text-generated images. The generated renderings are in line with the actual situation and the overall outline is clear. However, when verified on water meter data, the generated images appear a little blurry. The overall details are complete, and the size of the generated images cannot be changed at will. If the input image is too large, it will seriously affect the network operation speed. If the network structure is arbitrarily reduced, the quality of the generated images will be worse. Summary of the invention
[0005] The present invention proposes a method for generating high-resolution water meter images from text based on a diffusion model, which mainly includes two parts: diffusion-based model optimization and ESRGAN-based high-resolution network, and realizes the generation of high-resolution water meter images from text.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions.
[0007] A method for generating high-resolution images from text based on a diffusion model is designed based on U-net and ESRGAN networks and includes the following steps:
[0008] Step 1, data preparation: water meter image acquisition and water meter dataset preparation;
[0009] Step 2, model construction: The overall model structure is divided into two model architectures including:
[0010] Step 2.1, the first model architecture includes VAE autoencoder, CLIP model and diffusion denoising U-Net network;
[0011] Among them, the VAE autoencoder is used to embed the water meter image into a latent space. In the encoder part of the VAE, the water meter image is mapped to a latent space, and the decoder is responsible for reconstructing the image from the latent space. The key to VAE is to train by optimizing the variational lower bound (ELBO) to ensure that the generated image has good visual effects while retaining enough image information in the latent space;
[0012] CLIP model: The CLIP model combines image and text information, and can learn text and images in contrast, so that the model can generate images related to it through text prompts. In this architecture, CLIP embeds the text description of the water meter and the corresponding water meter image into a shared latent space to help incorporate the semantic information of the text when generating images.
[0013] Diffusion denoising U-Net network: This network architecture uses U-Net as the basic structure and combines the idea of diffusion denoising. The diffusion model gradually adds noise to the image, and through multiple iterations of the denoising process, a clear image is finally generated. The U-Net network includes the following modules:
[0014] Encoder: Extract image features through convolutional layers and gradually compress image information.
[0015] Middle block: Contains the core denoising module, which gradually removes noise and restores the structure of the image.
[0016] Decoder: gradually restores the features to the original image size and finally generates a clear image.
[0017] Step 2.2: The second model architecture adopts the ESRGAN network model architecture, including:
[0018] Generator: The generator is the core part of ESRGAN, responsible for generating high-resolution images from low-resolution images. It uses residual blocks to process image details and generate finer image features.
[0019] Discriminator: The discriminator is used to distinguish between real images and generated images. It guides the optimization of the generator by determining whether the input image comes from a real dataset. A convolutional neural network (CNN) is used to build the discriminator.
[0020] Adversarial loss: Adversarial loss trains the game between the generator and the discriminator, enabling the generator to generate high-resolution images that are increasingly close to real images.
[0021] Perceptual loss: Perceptual loss optimizes the visual quality of images by calculating the differences in images in a high-level feature space, making the generated images more realistic.
[0022] Pixel loss: The pixel-level loss function is used to ensure that the pixel accuracy of the generated image is close to the real image, usually calculated using the mean square error (MSE).
[0023] Step 3, model training: Divide the water meter image dataset into training and validation sets, ensure that the training set has sufficient diversity, and perform data augmentation (such as rotation, scaling, cropping, color transformation, etc.) to increase the robustness of the model. Use the optimizer to jointly train the VAE, CLIP, and diffusion models. For the ESRGAN model, use an adversarial training strategy, combining perceptual loss and pixel loss to improve image quality;
[0024] Step 4, model output: output the model with the best parameters after training to a *.pth file;
[0025] Step 5, model prediction: predict the trained .pth file, input text, and generate the corresponding image;
[0026] Step 6, super-resolution processing: Super-resolution processing is performed on the image generated by the text to make the generated image clearer and to allow the generated image to be enlarged or reduced to any size without losing accuracy.
[0027] In step 2.1 of the above-mentioned method for generating high-resolution water meter images from text based on a diffusion model, the present invention uses a multi-scale attention aggregation module as a bridge between the U-Net encoder and the decoder, which can capture features at different scales, helping to better capture detail features while retaining global semantic information and improving image generation quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0029] Figure 1 It is a framework diagram of the present invention.
[0030] Figure 2 The image super-resolution generation model architecture diagram of the present invention;
[0031] Figure 3 This is a data processing diagram of the multi-scale attention aggregation module of the present invention;
[0032] Figure 4 Results generated for the text of the present invention;
[0033] Figure 5 This is the result after super-resolution of the present invention; DETAILED DESCRIPTION
[0034] The present invention is mainly based on a diffusion model, and in order to further improve the clarity of the generated image, the present invention adds a super-resolution module on the basis of the diffusion model. The present invention can realize the clear processing of blurred images and generate images of various sizes.
[0035] Figure 1 This is a flowchart of the present invention, and the specific implementation of the present invention will be described below.
[0036] Step 1, data preparation: data collection and data set preparation;
[0037] Step 1.1, use the camera installed on the water meter to take a large number of photos of different types of water meters, and use random cropping and random reversal for cropping.
[0038] Step 1.2, convert the image information into hexadecimal, and then use the code to batch match the text description corresponding to each image.
[0039] Step 1.3: Divide the water meter images and their annotation information into a training set and a validation set, accounting for 80% and 20% respectively.
[0040] Step 2: According to the determined first model architecture, the original image is subjected to noise processing and the image text description is predicted to generate the expected image. The first loss function is determined according to the randomly added noise and the predicted removed noise, and the second model framework is trained to obtain the target second-stage image super-resolution generation model. The specific situation is described in the following steps:
[0041] Step 2.1, input the prepared data set into the first model architecture, use the encoder of the VAE autoencoder to encode the original image, and perform a random noise operation to convert it into a noise latent code, input the corresponding image text description into the pre-trained CLIP encoder to obtain text image features, and input the noise latent code and text image features into the diffusion denoising U-Net network for denoising. The text features in the diffusion denoising U-Net network are usually added to the image features through conditional activation functions, conditional convolutions, conditional noise, etc., and in the decoder stage, the text features help generate image content that is more in line with the description. Repeat the denoising operation, iterate T steps, and predict the image latent code from the diffusion denoising U-Net network, and then input the image latent code into the decoder of the VAE autoencoder to generate an image corresponding to the original image. The noise added each time in the random noise operation and the noise removed by each step of the iterative denoising operation are calculated for loss value, and the first loss function is determined, which is expressed as:
[0042]
[0043] Where z0 is the real image, t is the diffusion step size, z t is the noise image at time step t through the forward diffusion process, c is the text feature, is random noise, ∈ θ (z t ,t,c) is the noise predicted by the diffusion model.
[0044] Step 2.2, add the fixed-size and slightly blurred image generated by the first model architecture to the second model architecture. ESRGAN extracts the features of the low-resolution image through the generator, gradually upsamples and restores the details of the high-resolution image, and generates a clear and realistic super-resolution image through the comprehensive optimization of multiple loss functions (pixel loss, perceptual loss and adversarial loss). For blurred images, ESRGAN can effectively remove blur and restore lost details. ESRGAN combines multiple loss functions, including:
[0045] (1) Pixel loss:
[0046]
[0047] (2) Perceptual loss:
[0048]
[0049] (3) Fighting against losses:
[0050]
[0051] Step 3: For the designed diffusion model, use the training set to train the model and update the parameters. After every 20 rounds of training, save the model and observe the generated effect after saving. This process is iterated for 200 rounds. When the loss function remains stable and decreases discontinuously for multiple times, terminate the iteration and save the optimal parameter model to *.pth.
[0052] Step 4, model output: output the model with the best parameters after training to a *.pth file to facilitate subsequent prediction of the model;
[0053] Step 5: Use the trained optimal model to make predictions, input a desired text, and generate the corresponding image.
[0054] Step 6, finally, super-resolution processing is performed on the generated image. We use ESRGAN to achieve sharpness and arbitrary resizing of the image without losing details. Based on the previous SRGAN, ESRGAN removes all BN layers to enhance performance and reduce computational complexity. At the same time, we propose to replace the original basic block with a residual dense block (RRDB), which combines multi-layer residual networks and dense connections.
[0055] The above disclosure is only a specific embodiment of the present invention. According to the technical concept provided by the present invention, any changes that can be thought of by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A method for generating high-resolution images from text based on a diffusion model is designed based on U-net and ESRGAN networks and includes the following steps: Step 1, data preparation: water meter image acquisition and water meter dataset preparation; Step 2, model construction: The overall model structure is divided into two model architectures including: Step 2.1, the first model architecture includes VAE autoencoder, CLIP model and diffusion denoising U-Net network; Among them, the VAE autoencoder is used to embed the water meter image into a latent space. In the encoder part of the VAE, the water meter image is mapped to a latent space, and the decoder is responsible for reconstructing the image from the latent space. The key to VAE is to train by optimizing the variational lower bound (ELBO) to ensure that the generated image has good visual effects while retaining enough image information in the latent space; CLIP model: The CLIP model combines image and text information, and can learn text and images in contrast, so that the model can generate images related to it through text prompts. In this architecture, CLIP embeds the text description of the water meter and the corresponding water meter image into a shared latent space to help incorporate the semantic information of the text when generating images. Diffusion denoising U-Net network: This network architecture uses U-Net as the basic structure and combines the idea of diffusion denoising. The diffusion model gradually adds noise to the image, and through multiple iterations of the denoising process, a clear image is finally generated. The U-Net network includes the following modules: Encoder: Extract image features through convolutional layers and gradually compress image information. Middle block: Contains the core denoising module, which gradually removes noise and restores the structure of the image. Decoder: gradually restores the features to the original image size and finally generates a clear image. Step 2.2: The second model architecture adopts the ESRGAN network model architecture, including: Generator: The generator is the core part of ESRGAN, responsible for generating high-resolution images from low-resolution images. It uses residual blocks to process image details and generate finer image features. Discriminator: The discriminator is used to distinguish between real images and generated images. It guides the optimization of the generator by determining whether the input image comes from a real dataset. A convolutional neural network (CNN) is used to build the discriminator. Adversarial loss: Adversarial loss trains the game between the generator and the discriminator, enabling the generator to generate high-resolution images that are increasingly close to real images. Perceptual loss: Perceptual loss optimizes the visual quality of images by calculating the differences in images in a high-level feature space, making the generated images more realistic. Pixel loss: The pixel-level loss function is used to ensure that the pixel accuracy of the generated image is close to the real image, usually calculated using the mean square error (MSE). Step 3, model training: Divide the water meter image dataset into training and validation sets, ensure that the training set has sufficient diversity, and perform data augmentation (such as rotation, scaling, cropping, color transformation, etc.) to increase the robustness of the model. Use the optimizer to jointly train the VAE, CLIP, and diffusion models. For the ESRGAN model, use an adversarial training strategy, combining perceptual loss and pixel loss to improve image quality; Step 4, model output: output the model with the best parameters after training to a *.pth file; Step 5, model prediction: predict the trained .pth file, input text, and generate the corresponding image; Step 6, super-resolution processing: Super-resolution processing is performed on the image generated by the text to make the generated image clearer and to allow the generated image to be enlarged or reduced to any size without losing accuracy.