Stylized text graph diffusion model for intelligent water affairs and system implementation

By applying a text generation stylized image method based on diffusion model in water meter data recognition, the problem of low accuracy in existing models when identifying new types and environmental water meters is solved, high-quality and realistic water meter image generation is achieved, and recognition capabilities are improved.

CN120070639APending Publication Date: 2025-05-30BEIJING HONGCHENG XINDING INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205225.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When identifying water meter data, the existing deep network model has a single type of data set and a small amount of data, resulting in low robustness in the model, and the accuracy of water meter recognition in the face of new types and environments.

Method used

The text generation stylized image method based on diffusion model is adopted, and the stylized cultural image diffusion model and the cascaded diffusion model with external U-Net are improved, so the matching degree of text and image features is improved through a multi-layer cross-attention mechanism.

Benefits of technology

The generated water meter images are of high quality and fidelity, which can effectively meet the needs of smart water business, and improve the recognition ability of the model and the identification accuracy of new types and environmental water meters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070639A_ABST
    Figure CN120070639A_ABST
Patent Text Reader

Abstract

The invention provides a method for generating a water meter image based on a stylized text of a diffusion model, and aims to solve the problems of data set quantity shortage and single type during intelligent water meter reading recognition, an improved stylized text image diffusion model is used, and a water meter image with highly fused style and semantics is generated. And on the basis, a cascade diffusion model externally connected with U-Net is added to improve the quality of a final image, so that the problem of data set missing during intelligent reading identification is solved to a great extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent generation technology for generating images from text, and particularly to a generation network model that realizes stylization in the latent space suitable for multi-modal scenario environments. Specifically, it is a method for generating stylized images from text based on a diffusion model, belonging to the fields of computer vision and artificial intelligence generation. Background Art

[0002] In intelligent water services, water meter data is very important data. It is the basis for water service companies to accurately charge users and an important basis for pipe network maintenance and water pipe fault troubleshooting. Analyzing the water meter data of merchant factories can also effectively urge them to formulate reasonable water use plans to achieve water conservation. Traditional water meter data collection relies on manual meter reading, which requires a huge investment in manpower and material resources, and it is also difficult to guarantee the timeliness and accuracy of water meter data. With the increasing maturity of artificial intelligence technology and the rapid development of hardware technology, various identification technologies for water meter data have also been developed. For example, the object detection technology based on a deep network model is applied to the meter reading device, which can take pictures of water meter data in real time, upload it, and perform identification and calculation on the corresponding platform to obtain the water meter data. However, the deep network model needs real water meters to make a data set, send it into the deep network model for training, and learn different water meter features to achieve the goal of identifying water meter data. The accuracy of identification depends on the rationality of the network structure on the one hand, but more on the data set itself. In reality, there are a wide variety of water meter types, the water meter environment is complex and diverse, and the consumption of manpower and material resources for collecting training data results in a single type and small amount of data in the data set, leading to low model robustness and poor accuracy in identifying water meters of new types and environments.

[0003] In order to increase the diversity of the water meter data set and enhance the identification ability of the model, it is necessary to expand the types and quantities of the data set. Traditional data augmentation methods, such as image brightness adjustment, scaling, horizontal flipping, etc., only increase the data quantity, and the styles and environmental factors in the water meters are still single. The style transfer method based on deep learning can effectively increase the styles of water meter data, but it is difficult to customize and generate stylized water meters with specific features and difficult to generate more realistic water meter data. The text-to-image model based on the diffusion model can effectively generate specific water meter data, but the existing text-to-image models lack stronger controllability, and it is difficult to achieve the purpose of the styles and quality of the generated water meter image data.

[0004] In this case, there is an urgent need for a new network model for text-to-image generation to generate various types of higher-quality water meter image data. In the latest diffusion models, the integrated use of text-to-image models and stylization techniques is involved. However, when testing with a water meter dataset, two problems are found. First, the style features and text features cannot be well integrated, resulting in a strong sense of fragmentation in the generated stylized water meter images, which are very different from the actual water meter data. Second, the quality of the generated stylized water meter images is relatively low, making it difficult to meet the requirements of the intelligent water service business. Summary of the Invention

[0005] The present invention proposes a method for generating stylized water meter images from text based on a diffusion model, which realizes the generation of semantic stylized water meter images using text and style images.

[0006] To achieve the above object, the present invention adopts the following technical solutions.

[0007] It consists of two parts: an improved stylized text-to-image diffusion model and a cascaded diffusion model based on an external U-Net. The former solves the problem of weak semanticity in the pictures generated by existing stylized text-to-image methods, and the latter solves the problem of low quality of the generated stylized pictures. The design includes the following steps:

[0008] Step 1, data preparation: data collection; data augmentation; text template production; stylized image collection; text-image pair production;

[0009] Step 2, construction of text encoder and image encoder: The CLIP (Contrastive Language-Image Pre-Training) model is one of the most commonly used text encoders in current image-text generation tasks. Due to its pre-training on a large amount of image-text data, CLIP has strong robustness when processing different styles and types of image-text, and can achieve better results;

[0010] Text encoder: Select the CLIP model as the text encoder structure to convert the semantic prompt into an embedding (vector) that can be combined with the image, so as to generate pictures through semantic control. Its main advantage is that even if the input text is not very accurate, CLIP can correctly understand and generate diverse images without being overly interfered.

[0011] Image encoder: Select the CLIP model as the image encoder structure as a feature extractor. And cooperate with the Query Transformer (Q-Former) to generate a style embedding that can be combined with the image. Compared with traditional methods of extracting styles, this can extract the styles hidden in the image contour edges, and the generated images are more effective;

[0012] Step 3, Improve the design of the stylized text-to-image diffusion model: The existing diffusion models mainly consist of two key steps, namely the diffusion process and the denoising process;

[0013] The diffusion model generates clear images through a multi-step noise addition and denoising process. The diffusion process gradually adds noise to the input image and finally restores it to a clear image through the denoising step;

[0014] The denoising process gradually removes noise through conditional probability modeling and restores image details;

[0015] In order to fully integrate the style features and semantic features, a decoupled cross-attention mechanism is introduced. By using multiple independent cross-attention layers, the matching degree of text and image features is improved, and the fusion of style and semantics is enhanced.

[0016] Step 4, Improve the training of the stylized text-to-image diffusion model: Train the designed improved stylized text-to-image diffusion model to make the generated images have a deep fusion of style and semantics;

[0017] Step 5, Design of the cascaded diffusion model with an external U-Net: The role of the cascaded diffusion model is to remove noise and improve the detail quality of the previously generated images with good style and semantics but poor quality, thereby further improving the quality of the final images and ensuring the generation of high-quality water meter images. Use cross-entropy loss, perceptual loss, and L1 regression loss to optimize the model. The perceptual loss calculates the image difference through the feature maps of the VGG network, and the L1 loss ensures pixel-level accuracy;

[0018] Step 6, Train the cascaded diffusion model with an external U-Net: Train the cascaded diffusion model to ensure that the generated images have a high degree of stylization and excellent quality. Through iterative optimization, generate water meter images that meet the requirements of the water service;

[0019] Step 7, Model prediction: Perform prediction on the trained.pth file to generate stylized water meter images that meet the actual application requirements; Description of the Drawings

[0020] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0021] Figure 1 It is a flowchart of the present invention.

[0022] Figure 2 It is a model architecture diagram of the present invention.

[0023] Figure 3 It is a schematic diagram of the improved stylized text-to-image diffusion model of the present invention.

[0024] Figure 4Schematic diagram of the cascaded diffusion model with an external U-Net for the present invention. Detailed implementation manner

[0025] The present invention is mainly based on a diffusion model to generate pictures with highly integrated style semantics and high quality.

[0026] Figure 1 It is a flow block diagram of the present invention. The following will elaborate on the specific implementation of the present invention.

[0027] Step 1, data preparation:

[0028] Step 1.1, data collection. Collect a large number of image data related to water meters and equip each image with an English description. The image data should include different angles of water meters, different lighting conditions, different water meter categories, etc.;

[0029] Step 1.2, data augmentation. Perform data augmentation on the collected images. The specific operations include rotation, flipping, scaling, cropping, color adjustment, etc. to improve the generalization ability of the model, and then add black borders at the top and bottom to supplement it into a standard 320*320 pixel picture;

[0030] Step 1.3, text template making. Make a suitable text description template according to the characteristics of water meter images. Each image should correspond to an English description, and the description content includes information such as the contour features, surface features, and background features of the water meter image.

[0031] Step 1.4, making of text-image pairs. Each water meter image needs to be equipped with an English description, and pair each image with its corresponding text description to make text-image pairs for training.

[0032] Step 1.5, collection of stylized images. Collect stylized image data related to the style of water meter images to ensure the diversity and representativeness of the stylized images for subsequent stylization training.

[0033] Figure 2 It is the model architecture diagram of the present invention. The specific implementation method is as follows:

[0034] Step 2, the CLIP encoder can process text and image inputs. Use CLIP to process text respectively to obtain semantic information and process stylized images to obtain style information;

[0035] Step 2.1, in terms of text processing, first, the tokenizer converts the text into individual tokens or words. After tokenization, the text input is transformed into an embedding vector through nn.Embedding and then enters the text encoder with 12 encoding layers. Each encoding layer first normalizes through a layer_norm layer to help stabilize training and improve the convergence speed. Next, the self-attention layer calculates the relationship between each word and other words to generate a weighted embedding representation. Residual connections and re-normalization help alleviate the vanishing gradient problem.

[0036] Step 2.2, in terms of style image processing, after the normalization operation (the image size processing has been completed previously) to adapt to the model input requirements, it is fed into the image encoder. The image encoder is based on a convolutional neural network (ResNet) and is processed through multiple convolutional layers, and cooperates with the query transformer (Q-Former) to gradually extract the features of the image.

[0037] Step 3, improve the style-based text-to-image diffusion model to achieve a high degree of fusion of style features and text features. The model structure is as Figure 3 shown.

[0038] The core idea of the diffusion model is to add noise to the input image through a multi-step process to generate a noisy image, and finally restore the noisy image to a clear image through the denoising process. The mathematical model of the diffusion process is as follows:

[0039] q(x t ∣x t-1 ) = N(x t ; √(1 - β t )x t-1 , β t I)

[0040] where βt is the noise intensity, which increases as the number of time steps increases.

[0041] The denoising process gradually removes the noise through learning to restore a clear image. The denoising process of the model can be modeled by conditional probability:

[0042] p θ (x t-1 ∣x t ) = N(x t-1 ; μ θ (x t , t), Σθ(x t , t))

[0043] where μ θ (x t , t) is the denoising mean predicted by the neural network, and Σ θ (x t , t) is the covariance matrix.

[0044] To enhance the integration of style and semantics, the present invention introduces a decoupled cross-attention mechanism. The interaction between text features and image features is processed through multiple independent cross-attention layers to improve the semantic and style matching degree. The cross-attention calculation formula for text features and image features is as follows:

[0045]

[0046] where Q = ZW q , K = C t W k , V` = C t W v , and linear transformations are performed using different weight matrices.

[0047] After the text features and style features are processed by CLIP, they are fed into the decoupled cross-attention layer in the U-Net of the diffusion model to achieve a high degree of integration of style and semantics. The decoupled cross-attention mechanism uses multiple layers and activation functions. First, the text and style feature vectors are respectively transformed through independent linear layers (fully connected layers) to generate corresponding queries (Query), keys (Key), and values (Value). Among them, the dot product result of the query (Query) and the key (Key) is scaled and processed through the Softmax activation function to generate attention weights. To enhance the stability of training and reduce the problem of gradient disappearance, a residual connection is added after each attention operation, that is, the input features are directly added to the attention output. In addition, a layer normalization step follows each residual connection to help normalize the data distribution and improve the convergence speed of the model; there is a feed-forward network composed of fully connected layers behind each attention block. Here, it includes at least one hidden layer and the activation function ReLU (Rectified Linear Unit) for further processing features and increasing the non-linear expression ability;

[0048] Step 4, preliminarily train the improved style-based text-to-image diffusion model. If the generated stylized image does not match the semantics, it indicates that there is a problem with the design of the CLIP text encoder and it is necessary to return for inspection. If the style and semantics of the generated stylized image are not strongly integrated, it indicates that the decoupled attention layer of the improved style-based text-to-image diffusion model needs to be fine-tuned, modify it and retrain. If there is good integration, proceed to the next step of cascaded diffusion model training;

[0049] Step 5, the cascaded diffusion model with an external U-Net is used to improve the quality of stylized water meter images and refine the detailed parts in the images. By connecting the external U-Net structure, the denoising ability of the model is enhanced. The U-Net structure can transfer low-level features through its skip connections, thereby enhancing the detailed performance of the image.

[0050] To ensure that the images generated by the cascaded diffusion model have a high enough quality, it is necessary to reasonably design the loss function to guide the training process. The commonly used loss functions include:

[0051] Cross-entropy loss, mainly used for classification tasks, is usually used to train the discriminator in a generative adversarial network (GAN). It measures the difference in label classification between the generated image and the real image. However, in image generation tasks, cross-entropy loss is not commonly used as the only loss function. More often, it is used in combination with other losses.

[0052] Perceptual Loss, a commonly used loss function in deep learning image generation, measures the difference in the high-level feature space of images rather than at the pixel level. Its basic idea is to extract the feature representation of the image through a pre-trained convolutional neural network (such as VGG), and then calculate the difference between these feature representations: C i

[0053]

[0054] where Φ i represents the features extracted at the i-th layer, C i , H i , W i are the number of channels, height, and width of the feature map at this layer, and x and x′ represent the real image and the generated image respectively.

[0055] L1 regression loss is a common method for calculating the pixel-level error of images, especially suitable for detail restoration tasks. It constrains the image generation process by calculating the difference between the generated image and the real image at the pixel level. The L1 loss can effectively reduce the blurring phenomenon in the image and help the model reconstruct the image details more precisely:

[0056] L 1 = ||x - x′|| 1

[0057] where x is the target real image and x′ is the generated image. The L 1 norm calculates the sum of the absolute values of the pixel differences between the generated image and the real image, intuitively reflecting the difference between the generated image and the target image at the pixel level.

[0058] The above loss functions are combined with weights to form the total loss function:

[0059] L total = λ 1 L perceptual + λ 2 L 1 + λ 3 L cross-entropy

[0060] where λ 1 、λ 2 and λ 3 are weight coefficients. By fine-tuning the total loss function, the model can optimize the image details and maintain style consistency to improve the quality of the final water meter image.

[0061] Step 7: Conduct preliminary training on the cascaded diffusion model with an external U-Net, and the input is the stylized image generated by the improved style-based text-to-image diffusion model. If the generated result image does not have better quality or is even worse compared to the input image, it indicates that the various modules of the cascaded diffusion model need to be fine-tuned until higher-quality images can be generated.

[0062] Step 8: Train the model from scratch using the training set and update the parameters. After every 20 rounds of training, save the model once and observe the generated effect after saving. This process is iterated 200 times. When the loss function remains stable and does not continuously decrease multiple times, terminate the iteration and save the model with the optimal parameters to *.pth.

[0063] Step 9: Model output: Output the model with the optimal parameters after training to the *.pth file for subsequent prediction use of the model.

[0064] Step 10: Use the trained optimal model for prediction. Input a piece of text and a style image that you want to achieve, and generate the corresponding stylized water meter image.

[0065] The above-disclosed are only specific embodiments of the present invention. All changes that can be contemplated by those skilled in the art based on the technical idea provided by the present invention should fall within the protection scope of the present invention.

Claims

1. A two-part model based on an improved stylized text image diffusion model and a cascade diffusion model based on an external U-Net. The former solves the problem that the images generated by the existing stylized text image method are not semantically strong, and the latter solves the problem that the generated stylized images are of low quality. The design includes the following steps: Step 1, data preparation: data collection; Data augmentation; text template creation; stylized image collection; text-image pair creation; Step 2: Constructing text encoder and image encoder: The CLIP (Contrastive Language-Image Pre-Training) model is one of the most commonly used text encoders in image-text generation tasks. CLIP is pre-trained on a large amount of image-text data, so it has strong robustness in processing image-text of different styles and types, and can achieve better results. Text encoder: The CLIP model is selected as the text encoder structure to convert the semantic prompt into an embedding (vector) that can be combined with the image to generate images through semantic control. Its main advantage is that even if the input text is not very accurate, CLIP can correctly understand and generate diverse images without too much interference. Image encoder: The CLIP model is used as the image encoder structure and as a feature extractor. It is used in conjunction with the query converter (Q-Former) to generate style embeddings that can be combined with images. Compared with the traditional way of extracting style, this method can extract the style hidden in the edge of the image contour, and the generated image is more effective; Step 3: Improve the design of the diffusion model for stylized cultural images: The existing diffusion model mainly consists of two key steps, namely the diffusion process and the denoising process; The diffusion model generates a clear image through a multi-step denoising and denoising process. The diffusion process gradually adds noise to the input image and finally restores it to a clear image through the denoising step; The denoising process gradually removes noise and restores image details through conditional probability modeling; In order to fully integrate style features with semantic features, a decoupled cross-attention mechanism is introduced. Through multiple independent cross-attention layers, the matching degree of text and image features is improved, thereby enhancing the fusion of style and semantics. Step 4: Improved stylized text image diffusion model training: Train the designed improved stylized text image diffusion model so that the generated image has a deep fusion of style and semantics; Step 5, design of cascade diffusion model of external U-Net: The role of cascade diffusion model is to remove noise and improve detail quality of the pictures with good style and semantics but poor quality generated by the former, so as to further improve the quality of the final picture and ensure the generation of high-quality water meter images. Use cross entropy loss, perceptual loss and L1 regression loss to optimize the model. Perceptual loss calculates image differences through feature mapping of VGG network, and L1 loss ensures pixel-level accuracy. Step 6: Training of the cascade diffusion model of the external U-Net: Train the cascade diffusion model to ensure that the generated images are highly stylized and of high quality. Through iterative optimization, water meter images that meet water service requirements are generated; Step 7, model prediction: predict the trained .pth file to generate a stylized water meter image to meet actual application needs.