Water meter data enhancement method of any style migration model based on large convolution kernel
Through style transfer technology combined with large convolutional kernel networks, diversified water meter images are generated, which solves the problems of large manpower and material investment and poor environmental adaptability of existing water meter data acquisition methods, and improves data quality and model robustness.
Patent Information
- Application Number
- CN202510215958.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The existing water meter data collection methods require a lot of manpower and material investment, and it is difficult to collect complete water meter data in various environments, affecting the intelligent monitoring of the water system.
Using a water meter data enhancement method based on style migration, a more realistic and diverse water meter image is generated by combining the original water meter image with a good style reference style image, using the RepLKNet network and the VGG16 network of a large convolution kernel for feature extraction and fusion.
It effectively improves the quality and accuracy of water meter data, enhances the robustness and generalization capabilities of the model, reduces the resource consumption of data annotations, and is suitable for water meter data acquisition in various environments.
Smart Images

Figure CN120147110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a water meter data enhancement method based on style transfer, and in particular to a new technology for improving the quality and accuracy of water meter data. Specifically, it is a method that applies style transfer technology to combine the original water meter image with a reference style image with good style, so as to effectively enhance the water meter data. It belongs to the fields of computer vision and image generation Background Art
[0002] The object detection model is an effective method for water service data collection. However, due to the limitations of the object detection model, a large amount of data sets need to be collected for training the model; moreover, the annotation of water meter data also requires a large amount of manpower and material resources; in addition, it is difficult to collect all the water meter states in actual production and life, and it is impossible to complete the recognition of any water meter reading in life, which affects the intelligent monitoring of the water service system
[0003] With the rapid development of artificial intelligence in various fields, intelligent construction is gradually emerging in various industries. The object detection algorithm based on deep learning is an indispensable method in intelligent construction. In the current construction of intelligent water service, intelligent water meters can collect and upload water meter data at any time to realize the collection of water meter data. However, this device relies on power supply and is difficult to support operation in various environments; in addition, in some old rural areas and traditional communities, the construction of intelligent water meters is difficult and the installation cost is huge. Therefore, a camera transmission device installed on traditional water meters is needed. This device uploads the water meter image to the server side in real time by taking pictures, and uses an object detection network on the server side to recognize the reading of the water meter image. However, limited by the limitations of the water meter object detection network, a large amount of data sets are required for network training. The traditional acquisition method is to take a large number of images of different types of water meters and then manually annotate the images. However, it is difficult to collect all types of water meters on the market during data acquisition, and it is impossible to collect water meter images in various environments during the acquisition of water meter images. Therefore, the model trained according to these water meter images is not very robust; in addition, the annotation of the data set also consumes a large amount of resources, which has a great impact on the training of the object detection algorithm model
[0004] In this case, data augmentation is a crucial part of the water meter reading recognition task. Although traditional data augmentation methods can increase the diversity of training data to a certain extent, they face insurmountable defects, especially in the special context of processing water meter images. First of all, traditional augmentation methods often have difficulty simulating the changes of water meter images in different environments, such as water mist, water droplets, water scale, etc. These factors may lead to a large difference between the images after traditional augmentation and the real scene, thus affecting the generalization ability of the model. Secondly, traditional augmentation methods often rely on manually designed rules or simple transformations, and it is difficult to capture the high-level features and semantic information of images, which limits the performance of the model in complex scenarios. In view of the limitations of traditional augmentation methods, introducing style transfer as a data augmentation method has significant advantages. Style transfer can generate images with different styles and authenticity by learning the styles and features of water meter images, thus providing more diverse and more real data. This data augmentation method can not only simulate the changes of water meter images in different environments, but also expand the dataset of water meters, effectively improving the robustness and generalization ability of the model. In addition, the style transfer method can provide more diverse data without increasing the network complexity, avoiding the problem of decreased recognition efficiency caused by simply relying on augmentation in traditional augmentation methods. Therefore, compared with traditional data augmentation methods, the style transfer technology has greater potential and advantages in the water meter reading recognition task, can effectively improve the recognition accuracy, and is more suitable for actual application scenarios. Summary of the Invention
[0005] This paper proposes a water meter data augmentation method based on arbitrary style transfer with large convolutional kernels, which mainly includes a style extraction network based on large convolutional kernels and an image fusion algorithm based on whitening and coloring to achieve arbitrary style transfer of water meter images.
[0006] To achieve the above object, the present invention adopts the following design scheme:
[0007] A water meter data augmentation method based on arbitrary style transfer with large convolutional kernels is designed based on the VGG16 network and the RepLKNet network with large convolutional kernels, and includes the following steps:
[0008] Step 1, data preparation: data collection, which is divided into water meter images (also called content images) and style images;
[0009] Step 1.1, collect images of mainstream mechanical water meters on the market, and collect 100 style images on the network according to requirements, with a total of 5 categories of styles, 20 images for each style; including water mist, water droplets, water scale, dirty style and various real water meter image styles.
[0010] Step 2, Network Design: The content image and the style image are fused through the Whitening and Colorization (WCT) algorithm. Then, the content image and the style image are respectively used for feature extraction using the VGG16 network and the RepLKNet network. Finally, the extracted feature maps are fused again at the feature map level, and finally, the stylized image is output through the decoder;
[0011] Step 3, Model Optimization: The use of large convolutional kernels is accompanied by a quadratic increase in the number of parameters and FLOPs. Therefore, it is necessary to optimize the large convolutional kernel network, which includes the following sub-steps:
[0012] Step 3.1, Decompose the standard convolution into two steps: depthwise convolution and pointwise convolution;
[0013] Step 3.2, Depthwise Convolution is to perform independent convolution operations on each input channel of the input data. One convolutional kernel is responsible for one channel, and one channel is only convolved by one convolutional kernel. The number of channels of the feature map generated in this process is exactly the same as the number of input channels;
[0014] Step 3.3, Pointwise Convolution is to use a 1x1 convolutional kernel to perform convolution operations on each channel. Pointwise Convolution can be used to expand and compress the number of feature channels, and at the same time introduce non-linear transformations to make the network more expressive;
[0015] Step 3.4, By using Shortcut Connections, the RepLKNet network can capture local details without increasing additional parameters or the computational complexity of the network;
[0016] Step 3.5, Use the small kernel reparameterization method to make up for the optimization problem. First, construct a 3×3 layer parallel to the large layer, and then add the outputs after they pass through the Batch normalization (BN) layer. After training, merge the small kernel and BN parameters into the large kernel;
[0017] Step 4, Model Training: Train the designed style transfer network by sending the collected water meter content images and the style images collected from the network into the network for model training.
[0018] Step 5, Model Output: Output the trained model into a *.pth file for convenient subsequent testing;
[0019] Step 6, Model Verification: Load the model output in Step 5, input the water meter image and the style image, generate the stylized image, screen the available styles, input the target detection network, and set up a comparative experiment to verify the effectiveness of the algorithm for enhancing the water meter image;
[0020] In the above water meter data enhancement method based on arbitrary style transfer with large convolutional kernels, in Step 2, the whitening and coloring (WCT) algorithm is used to fuse the content image and the style image, initially extract the style features, and obtain the intermediate stylized image. Then, VGG16 and RepLKNet are used as the encoders of the content image and the style image respectively to obtain the feature maps of the water meter image and the content image. Then, calculations are performed through the intermediate stylized image, and the features of the stylized image are adjusted by backpropagation. Finally, perform WCT again to fuse the water meter image and the style image at the feature map level, and output the final stylized image after decoding;
[0021] In the above water meter data enhancement method based on arbitrary style transfer with large convolutional kernels, in Step 3.2, the computational amount is reduced to D k ×D k ×D F ×D F ×M; where (D k ,D k ,M) is the size of the feature map, (D F ,D F ,M) is the size of the convolutional kernel, and N is the number of convolutional kernels;
[0022] In the above water meter data enhancement method based on arbitrary style transfer with large convolutional kernels, in Step 3.3, the computational amount is reduced to D k ×D k ×N×M; so the computational amount of ordinary convolution is D k ×D k ×D F ×D F ×M×N, and the total computational amount of depthwise separable convolution is D k ×D k ×D F ×D F ×M + D k ×D k ×N×M;
[0023] In the above water meter data enhancement method based on arbitrary style transfer with large convolutional kernels, in Step 3.4, the design of the residual network is introduced. Cross-layer connection of features is achieved through the use of Shortcut Connections. By introducing the input of the previous level, the influence of the output change on the network weights is increased, and the gradient value of backpropagation also becomes larger, making the network training easier;
[0024] The above water meter data augmentation method for arbitrary style transfer based on large convolution kernels. In step 3.5, first calculate the mean μ of the input sample x i and variance β According to the calculated mean and variance, standardize the input to obtain Finally, by introducing learnable parameters γ and β for translation and scaling, restore the feature distribution that the original network is supposed to learn. Finally, by introducing learnable parameters γ and β for translation and scaling, restore the feature distribution that the original network is supposed to learn. Brief Description of the Drawings
[0025] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0026] Figure 1 It is the structural diagram of the present invention.
[0027] Figure 2 It is the structural diagram of the RepLKNet network.
[0028] Figure 3 It is the flowchart of the Stage module.
[0029] Figure 4 It is the comparison experiment diagram. Specific Embodiments
[0030] The present invention is mainly based on a style transfer model. In order to improve the robustness of the water meter target detection model from the data set preparation stage, the present invention provides a water meter data augmentation method for arbitrary style transfer based on large convolution kernels. The present invention can simulate the water meter field environment while retaining sufficient content information. On the premise of only changing the data augmentation algorithm, the accuracy of the SSD target detection algorithm is improved by 6.84%, and the accuracy of the YOLOv5 is improved by 6.56%.
[0031] Figure 1 It is the structural diagram of the present invention, and the specific implementation of the present invention will be described below.
[0032] Step 1, data preparation: data collection, which is divided into water meter images (also called content images) and style images;
[0033] Step 2, network design: fuse the content image and the style image through the whitening and coloring (WCT) algorithm, then extract the features of the content image and the style image using the VGG16 network and the RepLKNet network respectively, and finally perform feature fusion on the extracted feature maps at the feature map level, and finally output the stylized image through the decoder;
[0034] Step 3, Model Optimization: The use of large convolutional kernels is accompanied by a quadratic increase in the number of parameters and FLOPs. Therefore, it is necessary to optimize the large convolutional kernel network, which includes the following sub-steps:
[0035] Step 3.1, Decompose the standard convolution into two steps: depthwise convolution and pointwise convolution;
[0036] Step 3.2, Depthwise Convolution is to perform independent convolution operations on each input channel of the input data. One convolutional kernel is responsible for one channel, and only one convolutional kernel convolves one channel. The number of channels of the feature map generated in this process is exactly the same as the number of input channels;
[0037] Step 3.3, Pointwise Convolution is to use a 1x1 convolutional kernel to perform convolution operations on each channel. Pointwise Convolution can be used to expand and compress the number of feature channels, and at the same time introduce non-linear transformations to make the network more expressive;
[0038] Step 3.4, By using Shortcut Connections, the RepLKNet network can capture local details without increasing additional parameters or the computational complexity of the network;
[0039] Step 3.5, Use the small kernel reparameterization method to make up for the optimization problem. First, construct a 3×3 layer parallel to the large layer, and then add the outputs after they pass through the Batch normalization (BN) layer. After training, merge the small kernel and BN parameters into the large kernel;
[0040] Step 4, Model Training: Train the designed style transfer network by feeding the collected water meter content images and the style images collected from the network into the network for model training.
[0041] Step 5, Model Output: Output the trained model into a *.pth file for convenient subsequent testing;
[0042] Step 6, Model Verification: Load the model output in Step 5, input the water meter image and the style image to generate a stylized image, screen available styles, input them into the target detection network, and set up a comparative experiment to verify the effectiveness of the algorithm for enhancing the water meter image.
Claims
1. A water meter data enhancement method based on arbitrary style transfer with large convolution kernels is designed based on VGG16 network and RepLKNet network with large convolution kernels, including the following steps: Step 1, data preparation: data collection, divided into water meter images (also called content images) and style images; Step 1.1, collect images of mainstream mechanical water meters on the market, and collect 100 style images on the Internet according to needs, including 5 major styles, 20 images for each style; including water mist, water droplets, scale, dirty styles and various real water meter image styles. Step 2, network design: The content image and the style image are fused through the whitening and colorization (WCT) algorithm, and then the content image and the style image are respectively extracted using the VGG16 network and the RepLKNet network. Finally, the extracted feature maps are fused again at the feature map level, and finally the styled image is output through the decoder; Step 2.1: For the input water meter image Ic and style image Is, first use the WCT algorithm to extract the color features of the style and fuse them to generate the intermediate stylized image Ics, as shown in formula (1). Then input the water meter image and the intermediate stylized image Ics into the VGG16 network for feature extraction to obtain z_cs and z_c. Input the style image and the intermediate stylized image Ics into the RepLKNet large convolution kernel network for feature extraction to obtain r_cs and z_s. Calculate the content loss (formula (2)) and style loss (formula (3)) respectively. Finally, the feature map extracted from z_s is fused again through the WCT algorithm and decoded to obtain the final stylized image stylized (formula 4). Ics=WCT(Ic,Is) (1) Lc=SSIM(z_cs,z_c) (2) Ls=Gram_loss(r_cs,z_s) (3) Stylized=Decoder(WCT(z_c,z_s)) (4) Step 3: Model optimization: The use of large convolution kernels is accompanied by a secondary increase in the number of parameters and FLOPs, so it is necessary to optimize the large convolution kernel network, which includes the following sub-steps: Step 3.1, decompose the standard convolution into two steps: depth convolution and point-by-point convolution; Step 3.2, first perform an independent convolution operation on each input channel of the input data. One convolution kernel is responsible for one channel, and one channel is convolved by only one convolution kernel. The number of feature map channels generated by this process is exactly the same as the number of input channels. Step 3.3, use a 1x1 convolution kernel to perform convolution operations on each channel. Point-by-point convolution can be used to expand and compress the number of feature channels, while introducing nonlinear transformations to make the network more expressive; Step 3.4, by using Shortcut Connections, the RepLKNet network can capture local details without adding additional parameters or increasing the complexity of network calculation; Step 3.5, use the small kernel reparameterization method to remedy the optimization problem. First, build a 3×3 layer parallel to the large layer, and then add up their outputs after the batch normalization (BN) layer. After training, merge the small kernel and BN parameters into the large kernel; Step 4: Model training: Train the designed style transfer network, and send the collected water meter content images and style images collected on the Internet into the network for model training. Step 4.1: During model training, the input batch_size is set to 4, that is, 4 water meter images and content images are input at the same time each time, each batch_size is 1 iter, and the structural similarity (SSIM) between the stylized image and the content image is calculated every 10,000 iters. The total iter of the model is set to 150,000. If the model converges prematurely, the training is terminated and the current model is saved. Step 5: Model output: Output the trained model into a *.pth file for subsequent testing. Step 6, model verification: load the model output in step 5, input the water meter image and style image, generate a stylized image, filter the available styles, input the target detection network, and set up a comparative test to verify the effectiveness of the algorithm for water meter image enhancement; Step 6.1, load the trained model, and verify the effectiveness of the style transfer model by comparing the SSD and YOLOv5 target detection networks. The specific verification experiment process is shown in Figure 4; Step 6.2, send the original data to the target detection network for direct training; Step 6.3, use traditional data enhancement methods (a total of 12 methods), such as adding salt and pepper noise, Gaussian blur, detail enhancement, etc. For each image, randomly select n enhancement methods from the 12 methods (n is 3, 5, 7, 9, 11), and expand the annotated xml file, then add the original data to the data set, and get a total of (2000×n+2000) water meter images and corresponding annotation files, and finally send them to the target detection network for training; Step 6.4, for each image, n style images are randomly selected from 100 style images for enhancement; In step 6.5, each image is enhanced n times, and each time it is randomly selected whether to use the traditional enhancement method or the style transfer algorithm for enhancement.
2. A water meter data enhancement method based on arbitrary style transfer of large convolution kernels as claimed in claim 1, characterized in that In step 2, the content image and the style image are fused using the whitening and coloring (WCT) algorithm to preliminarily extract style features and obtain an intermediate stylized image. Then, VGG16 and RepLKNet are used as encoders for the content image and the style image, respectively, to obtain feature maps of the water meter image and the content image. The intermediate stylized image is then used for calculation and back propagation to adjust the features of the stylized image. Finally, WCT is performed again to fuse the water meter image and the style image at the feature map level, and the final stylized image is output after decoding.
3. The water meter data enhancement method based on arbitrary style transfer of large convolution kernels as claimed in claim 1, characterized in that In step 3.2, the amount of computation is reduced to D by depth convolution. k ×D k ×D F ×D F ×M; where (D k , D k , M) is the size of the feature map, (D F , D F , M) is the size of the convolution kernel, and N is the number of convolution kernels; The above-mentioned water meter data enhancement method based on arbitrary style transfer of large convolution kernel, in step 3.3, the amount of calculation is reduced to D by deep convolution k ×D k ×N×M; therefore, the amount of calculation for ordinary convolution is D k ×D k ×D F ×D F ×M×N, and the total computational effort of the depthwise separable convolution is D k ×D k ×D F ×D F ×M+D k ×D k ×N×M; 4. The water meter data enhancement method based on arbitrary style transfer of large convolution kernels as claimed in claim 1, characterized in that In step 3.4, the residual network design is introduced. By using ShortcutConnections to realize cross-layer links of features, the input of the previous layer is introduced to increase the impact of output changes on network weights, and the gradient value of back propagation also becomes larger, making network training easier; 5. The water meter data enhancement method based on arbitrary style transfer of large convolution kernels as claimed in claim 1, characterized in that In step 3.5, first calculate the input sample x i The mean μ β With variance The input is standardized according to the calculated mean and variance. Finally, by introducing learnable parameters γ and β for translation and scaling, the feature distribution that the original network wants to learn is restored.