Configuration-free edge water meter water quantity identification and data quality control method

Through data set expansion, model network optimization and self-attention time series interpolation model, the diversity and anti-interference problems of water meter recognition technology are solved, the accuracy and data integrity of water meter recognition are improved, and the management efficiency and resource allocation of water systems are improved.

CN120299010APending Publication Date: 2025-07-11BEIJING HONGCHENG XINDING INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385659.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing water meter identification technology has significant limitations in terms of diversity, anti-interference and low light adaptability, which affects the accuracy and efficiency of identification. At the same time, the data integrity and accuracy problems in water systems need to be solved urgently, resulting in uneven resource allocation and insufficient water supply.

Method used

Data set expansion and enhancement, model network optimization design, preprocessing technology research and identification model multi-platform deployment is adopted. Through technologies such as image enhancement, network structure modification, feature fusion, deep separable convolution and model pruning, data governance is carried out in combination with the self-attention time series interpolation model to improve identification accuracy and data integrity.

Benefits of technology

It improves the accuracy and adaptability of water meter identification, ensures the continuity and integrity of data, improves the decision-making efficiency and rationality of resource allocation in the water management system, and reduces the impact of data loss on water bill calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299010A_ABST
    Figure CN120299010A_ABST
Patent Text Reader

Abstract

The invention discloses a configuration-free edge water meter water quantity identification and data quality control method, and aims to solve the problems that the current domestic water quantity image identification accuracy cannot meet the requirement of a management department for resident or non-resident water quantity remote transmission, and the scale application of an image water quantity identification technology is limited due to insufficient identification precision. A water volume image identification optimization model and a water volume data governance model algorithm are formed through identification model optimization key technology research and data governance model algorithm research, so that the identification algorithm can be simultaneously suitable for mechanical instruments which realize display through numbers and pointers, the identification precision reaches 99% or above, and the identification time is less than 1 second; the algorithm accuracy of the water volume data management model reaches 90% or above, the interpolation time of the model is controlled at the second level, and the processing requirement of large-scale water volume data can be responded in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent detection and recognition method for traditional water meter measurement, specifically a method for identifying water volume of edge water meters without configuration and controlling data quality, belonging to the fields of computer vision and pattern recognition. Background Art

[0002] With the acceleration of the urbanization process, water resource management has become increasingly important in modern society. As the core equipment for water resource measurement and management, water meters are widely used in various scenarios such as households, industries, and commerce. With the development of intelligent technologies, water meter recognition technology has gradually shifted from manual meter reading to automatic recognition. However, the challenges faced in the automatic recognition process are very complex, especially in aspects such as the variety of water meter types, different dial styles, and complex on-site environments, making the accurate recognition of water meters a difficult task. Existing water meter reading recognition algorithms still have significant limitations in terms of diversity, anti-interference ability, and low-light adaptability. These problems significantly affect the accuracy and efficiency of water meter recognition in practical applications. Therefore, it is necessary to further optimize the algorithms, introduce more intelligent processing mechanisms, and fully consider various complex usage scenarios during the design to address the deficiencies of current technologies.

[0003] In addition to the challenges of water meter recognition technology, the issues of data integrity and accuracy in the water service system also need to be urgently solved. Any form of data loss, data anomaly, or data delay will directly affect the decision-making efficiency and resource allocation rationality of the water service management system. In the daily operation of the water service system, the integrity of water volume data directly affects resource scheduling and water supply optimization. The loss or error of water meter data may cause the management department to be unable to accurately evaluate the water usage situation in a certain area, thereby leading to uneven resource allocation or insufficient water supply. Especially during peak water usage periods, if accurate water usage data cannot be obtained, the water service department may have difficulty meeting the water demand in this area, thereby affecting the normal life of users. In addition, data loss will also affect the accuracy of water fee calculation, resulting in errors in the charging system, thereby affecting the revenue management of water service enterprises. Therefore, ensuring the accuracy and continuity of water meter data is crucial for improving daily operation efficiency.

[0004] To solve the above problems and improve the accuracy and adaptability of water meter recognition, it is urgently necessary to optimize the existing water meter recognition technology. First of all, it is necessary to develop more general recognition algorithms that can achieve efficient and accurate recognition on different types and styles of water meters. At the same time, the image processing algorithm should have stronger anti-interference ability and be able to effectively handle recognition problems in complex environments such as water mist, water bubbles, and dirt. In addition, by introducing emerging technologies such as deep learning and artificial intelligence, a more intelligent water volume image recognition model is established to further improve the adaptability and recognition speed of the system.

[0005] In addition, with the rapid development of machine learning and big data technologies, the water volume data governance model has become an effective tool for dealing with complex data environments. Traditional interpolation methods such as linear interpolation and spline interpolation show limitations when facing large-scale data missing and non-linear data, while the imputation model based on deep learning can better handle complex, heterogeneous and large-scale data environments. Through these advanced algorithms, the water management system can still maintain a high decision-making accuracy even when the data collection is incomplete.

[0006] To solve the deficiencies of the existing technologies, the present invention provides dataset expansion and enhancement, model network optimization design, preprocessing technology research, and recognition model multi-platform deployment.

[0007] To achieve the above object, the present invention adopts the following technical solutions.

[0008] Step 1 Dataset expansion and enhancement:

[0009] Step 1.1 Collect 20,000 water meter images under different environmental conditions and at different time periods from water meters of different manufacturers and different types, evenly divide them into a training set and a validation set, and label the positions and readings of the numbers and pointers of each image.

[0010] Step 1.2 Perform image enhancement using traditional image enhancement based on geometric transformation, traditional image enhancement based on color transformation, traditional image enhancement based on mixed data, and traditional image enhancement algorithm based on OpenCV;

[0011] Step 1.3 Use any style transfer technology to separate the "content" of one image and the "style" of another image, and then apply different styles to the content image.

[0012] Step 1.4 Regularly collect new data.

[0013] Step 2 Model network optimization design to improve the model recognition accuracy:

[0014] Step 2.1 Modify the network structure and add a network attention module;

[0015] Step 2.2 Use a feature pyramid network for multi-scale feature fusion; in the network structure, compress the feature map through downsampling operations (such as pooling layers); restore the resolution of the feature map through upsampling operations (such as deconvolution layers), and fuse high-level semantic features with low-level detail features.

[0016] Step 2.3 Detect the same target at different levels of the network and fuse the feature maps from different levels

[0017] Step 2.4 Set multiple detection heads on different feature layers of the model to perform object detection on the feature maps of each layer.

[0018] Step 2.4 Fuse multiple feature maps at different levels by methods such as weighted summation and concatenation, and then perform object detection on the fused feature map.

[0019] Step 3 Optimize the network model to accelerate the model recognition speed:

[0020] Step 3.1 Use depthwise separable convolutions to reduce the number of parameters and computational complexity.

[0021] Step 3.2 Reduce the floating-point weights and activation function values of the model to 8-bit or lower representations (such as integers) to reduce memory occupancy and computational complexity.

[0022] Step 3.3 Use model pruning to remove neurons, channels, or connections that contribute less to the model performance, reducing the model width.

[0023] Step 3.4 Use model knowledge distillation to reduce the model depth and width.

[0024] Step 4 Restoration of compressed images:

[0025] Step 4.1 Feature extraction and restoration: Use a convolutional neural network to extract key features from the compressed image and compare them with the features of the original high-definition image to identify the features lost or damaged during the compression process.

[0026] Step 4.2 Detail reconstruction: Learn the differences between the original image and the compressed image, and the model generates new detail information to supplement or replace the details lost during the compression process, improving the visual quality of the compressed image.

[0027] Step 4.3 Optimize the feature distribution: Optimize the feature distribution of the compressed image by methods such as the maximum mean discrepancy (MMD).

[0028] Step 5 Build and train a water volume data governance model:

[0029] Step 5.1 Use techniques such as time shift, mixup augmentation, scaling, noise, and perturbation to perform data augmentation on historical water usage data.

[0030] Step 5.2 After data augmentation, use one-hot encoding to encode different types of water usage data.

[0031] Step 5.3 The encoded data is input into a self-attention based time series imputation model for training.

[0032] Step 5.3.1 Divide the dataset into a training set, a validation set, and a test set. The proportion of the dataset division is 70% for the training set, 15% for the validation set, and 15% for the test set;

[0033] Step 5.3.2 Use the Bayesian optimization algorithm to optimize the network model. By constructing a probability model of hyperparameters and selecting the next hyperparameter combination to be evaluated based on the output results of the model, the optimal solution is gradually approximated.

[0034] Step 5.3.3 Comprehensively use step decay, exponential decay, and an adaptive learning rate optimizer to automatically adjust the learning rate at different stages of training.

[0035] Step 5.3.4 Adopt an early stopping strategy. When the performance of the validation set does not improve for 10 consecutive epochs, the model training stops, and the model parameters with the best performance on the validation set are retained.

[0036] Step 5.4 Evaluate and verify the recognition accuracy, generalization ability, and running efficiency of the model:

[0037] Regarding the recognition accuracy of the model, the imputation accuracy of the model is quantified through three types of error metrics: root mean square error, mean absolute error, and coefficient of determination.

[0038] Step 5.4.2 Test the model under different missing rates, including the above three types of error metrics under data missing rates such as 5%, 10%, 20%, 30%, etc.;

[0039] Step 5.4.3 Use the water system datasets from different cities, seasons, and climate conditions to evaluate the generalization ability of the model.

[0040] Step 5.4.4 Use the imputation time of the model, the running efficiency on a single standard server (with 32GB of memory and CPU usage not exceeding 50%), the GPU acceleration efficiency, and the model response time to verify the running efficiency of the model;

[0041] Step 5.5 Result optimization: After being processed by the neural network model, the system will output various results. First, the model marks the detected abnormal data. Second, the model cleans and imputes the missing data, filling the data gaps through prediction. Finally, the model provides a prediction of future water volumes.

[0042] Figure 1 Flowchart of dataset augmentation and enhancement

[0043] Figure 2 Flowchart of network model optimization customization

[0044] Figure 3 Flowchart of water volume imputation

[0045] Figure 4 Schematic diagram of data augmentation based on time offset

[0046] Figure 5 Schematic diagram of data augmentation based on hybrid augmentation

[0047] Figure 6 Schematic diagram of data augmentation based on noise and perturbation

[0048] Figure 7 Schematic diagram of data augmentation based on scaling

[0049] Figure 8 Water volume data governance model based on SAITS

[0050] Time series of water meter A readings before model processing in Figure 9(a)

[0051] Time series of water meter A readings after model processing in Figure 9(b)

[0052] Time series of water meter B readings before model processing in Figure 10(a)

[0053] Time series of water meter B readings after model processing in Figure 10(b)

[0054] Figure 11 Before and after imputation of daily water consumption of water meter C

[0055] Time series of water meter D readings before model processing in Figure 12(a)

[0056] Time series of water meter D readings after model processing in Figure 12(b)

[0057] Specific implementation:

[0058] Step 1 The expansion and enhancement of the dataset are as Figure 1 shown:

[0059] Step 1.1 For water meters from different manufacturers and of different types, 20,000 water meter images are collected under different environmental conditions and at different time periods, and are evenly divided into a training set and a validation set. The positions and readings of the numbers and pointers of each image are labeled.

[0060] Step 1.2 Traditional image enhancement based on geometric transformation, traditional image enhancement based on color transformation, traditional image enhancement based on hybrid data, and traditional image enhancement algorithms based on OpenCV are used for image enhancement;

[0061] Step 1.3 Use any style transfer technology to separate the "content" of one image and the "style" of another image, and then apply different styles to the content image.

[0062] Step 1.4 Regularly collect new data.

[0063] Step 2 is as follows Figure 2 Optimize the network model as shown to improve the recognition accuracy and speed of the model:

[0064] Step 2.1 Modify the network structure and add a network attention module;

[0065] Step 2.2 Use a feature pyramid network for multi-scale feature fusion; in the network structure, compress the feature map through downsampling operations (such as pooling layers); restore the resolution of the feature map through upsampling operations (such as transposed convolution layers), and fuse high-level semantic features with low-level detail features.

[0066] Step 2.3 Detect the same target at different levels of the network and fuse the feature maps from different levels.

[0067] Step 2.4 Set multiple detection heads on different feature layers of the model to perform object detection on the feature map of each layer;

[0068] Step 2.4 Fuse multiple feature maps of different levels by weighted summation, concatenation, etc., and then perform object detection on the fused feature map;

[0069] Step 2.5 Use depthwise separable convolution to reduce the number of parameters and computational complexity;

[0070] Step 2.6 Reduce the floating-point weights and activation function values of the model to 8-bit or lower representation forms (such as integers) to reduce memory occupancy and computational complexity;

[0071] Step 2.7 Use model pruning to remove neurons, channels, or connections that contribute less to the model performance and reduce the model width;

[0072] Step 2.8 Use model knowledge distillation to reduce the model depth and width;

[0073] Step 3 Image preprocessing includes the restoration of compressed images:

[0074] Step 3.1 Feature extraction and restoration: Use a convolutional neural network to extract key features from the compressed image and compare them with the features of the original high-definition image to identify the features lost or damaged during the compression process.

[0075] Step 3.2 Detail reconstruction: Learn the differences between the original image and the compressed image, and the model generates new detail information to supplement or replace the details lost during the compression process to improve the visual quality of the compressed image.

[0076] Step 3.3 Optimize the feature distribution: Optimize the feature distribution of the compressed image through methods such as maximum mean discrepancy (MMD).

[0077] Step 4 is as follows Figure 3As shown in the figure, a water volume data governance model is constructed and trained:

[0078] Step 4.1: Use the time shift algorithm to simulate and generate data with different time series for historical water consumption data, use the hybrid enhancement model to simulate and generate data with diverse features, use noise and perturbation to simulate and generate noisy data in the real scenario, and use the scaling model to simulate and generate data with different scales for data enhancement.

[0079] The time shift technique generates multiple different time series through the transformation of time points, enabling the model to encounter more diverse time patterns during the training process, thereby improving the model's performance when facing unknown data.

[0080] Specifically, the time shift performs a mathematical transformation on each time point in the original time series through the polynomial deformation function shown in Equation (1). Let a certain time point be t, and the parameter controlling the degree of deformation be α. Then, the new time series can be generated through the following non-linear transformation function:

[0081] f(t) = t + α × t 2 (1)

[0082] In this formula, t represents the time point in the original time series, and α is the deformation coefficient used to control the degree of time shift. Through this function, each time point t in the original time series will be transformed into a new time point f(t), thereby generating a completely new series after time shift.

[0083] This non-linear transformation can "bend" or "stretch" the time axis to a certain extent, causing the position of the time points to shift. By adjusting the magnitude of the parameter α, the amplitude of the time shift can be flexibly controlled. When α is small, the amplitude of the shift is small, and the difference between the new time series and the original time series is small, mainly reflected in the slight adjustment of the time points. In this case, the time structure of the data remains basically unchanged, but slight perturbations are introduced, providing the model with slightly different time series samples.

[0084] When α is large, the amplitude of the shift increases significantly, and the time points of the time series may be adjusted greatly, resulting in a large change in the local structure of the time series. In this case, the generated new time series, while maintaining the trend of the original data, has a more diverse time arrangement, effectively enriching the structural diversity of the training dataset.

[0085] The hybrid enhancement technique fuses the features in different time series to generate a new series with more diverse features. This new series not only retains the useful information in the original series but also introduces new data variation forms through randomized mixing ratios.

[0086] In the hybrid enhancement method, usually two (or more) time series as shown in Figure 3 are selected for combination.

[0087] Suppose the two time series are X1 and X2 respectively. To generate a new time series X′, it is first necessary to determine their mixing ratio. This mixing ratio λ is randomly generated from the Beta distribution, and the value of λ is between 0 and 1. The Beta distribution is a distribution form commonly used in probability statistics and can flexibly control the distribution form of λ. Specifically, the generation process of λ can be expressed as:

[0088] λ ∼ Beta(α, α)

[0089] In this formula, α is the shape parameter of the Beta distribution, usually set between 0.2 and 0.4. This parameter determines the value distribution form of λ. When α is small, the value of λ is closer to 0 or 1, that is, one of the two time series X1 and X2 will dominate. When α is large, the value of λ is closer to 0.5, indicating that the generated time series X′ mixes the information in the two time series more evenly.

[0090] After determining the mixing ratio λ, the newly generated time series X′ can be obtained by linear combination. The specific formula is:

[0091] X' = λ × X1 + (1 - λ) × X2

[0092] In this formula, X′ is the new time series, X1 and X2 are the two original time series respectively, and λ is the mixing ratio generated above. Through this linear combination formula, X1′ contains both part of the information of the time series X1 and combines part of the characteristics of the time series X2, thus generating a new and diverse data series.

[0093] Simulate the data fluctuations in the real environment and improve the model's tolerance to noise. Additive noise and perturbations are introduced to the time series data as Figure 5 shown.

[0094] For each original time series data point Xt, an enhanced data point Xt′ is generated by adding a randomly generated noise term ∈t. The commonly used noise generation method is to sample from the normal distribution. The generation process of the noise term ∈t can be expressed as:

[0095] ∈ t ∼ N(μ, σ 2 )

[0096] Among them, N(μ,σ2) represents a normal distribution with a mean of μ and a variance of σ2. The mean μ of the noise term determines the central position of the noise, usually set to 0 to ensure that the noise does not introduce systematic bias. The standard deviation σ of the noise, on the other hand, controls the intensity or amplitude of the noise and determines the range of fluctuations of the noise values.

[0097] The formula for generating the new data point Xt′ is as follows:

[0098] X' t = X t + ∈ t

[0099] Through the above formula, an enhanced data point Xt′ with noise can be generated for each time point t. In this process, the basic pattern Xt of the original time series is retained, and at the same time, randomness and perturbations are introduced by adding the noise term ∈t.

[0100] As Figure 6 shown, scaling generates data samples of different scales by adjusting the proportion of the original data.

[0101] Specifically, the implementation of the scaling technique is by applying a scaling factor δ to each data point in the time series. This scaling factor is usually randomly generated within a preset range, representing the proportional amplification or reduction of the original data. The set range of the scaling factor is [0.8, 1.2], meaning that during the data augmentation process, the amplitude of the data can vary between 80% and 120% of the original value.

[0102] Multiply the generated scaling factor δ point by point with each data point Xt in the original time series to obtain a new enhanced data sequence Xt′. The specific operation is as follows:

[0103] X′ t = δ × X t

[0104] Through this process, the value of each data point is adjusted to δ times the original value, thus generating a new time series.

[0105] Step 4.2 After data augmentation, use one-hot encoding to encode different types of water usage data;

[0106] One-hot encoding can independently identify different water usage categories, enabling the model to directly use these encodings for the water volume interpolation training of different water usage categories.

[0107] For the set C = {c1, c2,..., c 41} of 41 water usage categories defined in the urban water usage classification standard, each category c iAll can be transformed into a vector of length 41 through one-hot encoding The specific form is as follows:

[0108] v i =(v1, v2, …, v 41 )

[0109] Among them, the value of each element vj in the vector v i is determined by the following rules:

[0110]

[0111] The basic principle of this encoding method is that in the vector v i , the i-th element is set to 1, indicating the existence of the category ci, and the elements in other positions are all set to 0, indicating other categories outside this category. In this case, all categories ci have their unique one-hot encoding forms, thus ensuring the mutual independence between different categories.

[0112] Step 4.3 The encoded data is input into a self-attention-based time series imputation model (Self-Attention-Based Imputation for Time Series, SAITS) for training;

[0113] The self-attention-based time series imputation model is a neural network architecture for imputing missing values in time series, and its feature lies in introducing a diagonal mask in the self-attention mechanism. Specifically, the diagonal elements of the attention matrix of the self-attention mechanism are set to -∞ (set to -1×109 in actual applications), so that after passing through the softmax function, the attention weights on the diagonal will approach 0. The calculation method is as follows:

[0114]

[0115] This method relies on the weighted combination of two diagonally masked self-attention (Diagonally-Masked Self-Attention, DMSA) blocks, as Figure 5 shown. The DMSA block can explicitly capture the temporal dependence of the time series and the correlation between features, which is very crucial for improving the imputation accuracy and training speed. At the same time, the weighted combination design enables SAITS to dynamically assign weights to the learned representations from the two DMSA blocks according to the attention map and missing information. In this way, SAITS can more effectively handle the missing value problem in time series data. According to the goal, the model structure is adjusted so that it can successively output anomaly data markers, data imputation results, and future water volume predictions.

[0116] Step 4.3.1 Divide the dataset into a training set, a validation set, and a test set. The ratio of dataset division is 70% for the training set, 15% for the validation set, and 15% for the test set;

[0117] Step 4.3.2 Use the Bayesian optimization algorithm to optimize the network model. By constructing a probability model of hyperparameters and based on the output results of the model, select the next hyperparameter combination to be evaluated, so as to gradually approach the optimal solution.

[0118] Step 4.3.3 Comprehensively use step decay, exponential decay, and an adaptive learning rate optimizer to automatically adjust the learning rate at different stages of training.

[0119] Step 4.3.4 Adopt an early stopping strategy. When the performance of the validation set does not improve for 10 consecutive epochs, stop the model training and retain the model parameters with the best performance on the validation set.

[0120] Step 4.4 Evaluate and verify the recognition accuracy, generalization ability, and running efficiency of the model:

[0121] Regarding the recognition accuracy of the model, quantify the imputation accuracy of the model through three error metrics: root mean square error, mean absolute error, and coefficient of determination.

[0122] Root mean square error (RMSE) is a commonly used metric to measure the difference between the estimated values and the true values of the model. The smaller the value of RMSE, the higher the accuracy of the model and the smaller the error. Its calculation formula is:

[0123]

[0124] where n is the number of samples, yi is the true value of the i-th sample, is the estimated value of the i-th sample.

[0125] Mean absolute error (MAE) is another commonly used evaluation metric. Its calculation formula is:

[0126]

[0127] MAE measures the average error of the model by taking the average of the absolute values of all errors. Different from RMSE, MAE does not give extra weight to large errors, so it is less affected by outliers and is a more "gentle" way to measure errors.

[0128] Coefficient of determination (R 2 ) is also known as the determination coefficient and is an important metric to measure the fitting effect of the model. Its formula is:

[0129]

[0130] R 2 ranges from 0 to 1, where 1 represents a perfect fit, that is, the estimated value of the model exactly matches the true value; 0 means that the imputation effect of the model is the same as simply using the average value as the estimation result. R 2 can also take negative values, which indicates that the model performance is poor. R 2 Intuitively reflects the ability of the model to explain the changes in real data. A higher R 2 value indicates that the model can handle the imputation task well, while a lower R 2 may mean that the model is underfitting, or there is significant noise in the data.

[0131] Step 4.4.2 Test the model under different missing rates, including the above three types of error indicators at data missing rates such as 5%, 10%, 20%, 30%, etc.;

[0132] Step 4.4.3 Use the water system dataset from different cities, seasons, and climate conditions to evaluate the generalization ability of the model.

[0133] Step 4.4.4 Verify the running efficiency of the model by using the imputation time of the model, the running efficiency on a single standard server (with 32GB of memory and CPU usage not exceeding 50%), the GPU acceleration efficiency, and the model response time;

[0134] Step 4.5 Result optimization: After being processed by the neural network model, the system will output various results. First, the model marks the detected abnormal data. Second, the model cleans and imputes the missing data, filling the data gaps through prediction. Finally, the model provides a prediction of future water volume.

[0135] For extreme outliers, this kind of abnormal reading deviates greatly from the actual reading, which is caused by errors in the data collection process, equipment failures, accidental events, etc. Such anomalies have obvious abnormal characteristics, and the model recognition rate reaches 100%, as shown in Figure 9.

[0136] For marginal outliers, this kind of abnormal reading is close to the normal range of the data but slightly deviated. As shown in Figure 10, the gap between some outliers and the true value is even within 1 ton. Such anomalies are difficult to detect globally, and they belong to the small noise in the data and can only be captured in local data. The model recognition rate for this kind of anomaly is 95%.

[0137] The model can fill the missing values in the data to ensure the integrity and coherence of the data, and at the same time ensure the accuracy of the imputed values, making them close to the real data. The following is the imputation effect diagram. Figure 11 It is a bar chart of daily water consumption. The value on May 7 is the imputed value, which is close to the real water consumption.

[0138] Figure 12 shows the time series graph of water meter readings before and after interpolation. It can be seen that there are many missing values in the original data. After interpolation by the model, the curve becomes complete and continuous, and at the same time, it maintains the original trend, and the interpolation effect is remarkable.

[0139] By extracting some water meters with complete data, part of the data is manually removed to evaluate the interpolation effect of the model. After multiple tests, it is found that the deviation between the interpolated value of the model and the true value is small, and most of the deviation ranges between 1% and 15%.

Claims

1. A method for identifying water volume of edge water meters without configuration and controlling data quality, characterized in that, The steps are as follows: Step 1 Dataset augmentation and enhancement: Step 1.1 Collect 20,000 water meter images under different environmental conditions and at different time periods from water meters of different manufacturers and types, and evenly divide them into a training set and a validation set. Mark the positions and readings of the numbers and pointers in each image. Step 1.2 Use traditional image enhancement based on geometric transformation, traditional image enhancement based on color transformation, traditional image enhancement based on mixed data, and traditional image enhancement algorithms based on OpenCV for image enhancement; Step 1.3 Use any style transfer technique to separate the "content" of one image from the "style" of another image, and then apply different styles to the content image. Step 1.4 Regularly collect new data. Step 2 Optimize the model network design to improve the model recognition accuracy: Step 2.1 Modify the network structure and add a network attention module; Step 2.2 Use a feature pyramid network for multi-scale feature fusion; in the network structure, compress the feature map through downsampling operations (such as pooling layers); restore the resolution of the feature map through upsampling operations (such as deconvolution layers), and fuse high-level semantic features with low-level detail features. Step 2.3 Detect the same target at different levels of the network and fuse the feature maps from different levels Step 2.4 Set multiple detection heads (Detection Head) on different feature layers of the model to perform object detection on the feature map of each layer Step 2.4 Fuse multiple feature maps at different levels by means of weighted summation, concatenation, etc., and then perform object detection on the fused feature map; Step 3 Optimize the network model to speed up the model recognition speed: Step 3.1 Use depthwise separable convolution to reduce the number of parameters and computational complexity; Step 3.2 Reduce the floating-point weights and activation function values of the model to 8-bit or lower representation forms (such as integers) to reduce memory occupancy and computational complexity; Step 3.3 Use model pruning to remove neurons, channels or connections that contribute less to the model performance and reduce the model width; Step 3.4 Use model knowledge distillation to reduce the model depth and width; Step 4 Compressed image restoration: Step 4.1 Feature extraction and restoration: Use a convolutional neural network to extract key features from the compressed image and compare them with the features of the original high-definition image to identify the features lost or damaged during the compression process. Step 4.2 Detail reconstruction: Learn the differences between the original image and the compressed image, and the model generates new detail information to supplement or replace the details lost during the compression process to improve the visual quality of the compressed image. Step 4.3 Optimize the feature distribution: Optimize the feature distribution of the compressed image through methods such as maximum mean discrepancy (MMD). Step 5 Build and train a water volume data governance model: Step 5.1 Use techniques such as time shift, mixed enhancement, scaling, noise and perturbation to perform data enhancement on historical water consumption data; Step 5.2 After data enhancement, use one-hot encoding to encode different types of water consumption data; Step 5.3 The encoded data is input into a self-attention-based time series imputation model for training; Step 5.3.1 The dataset is divided into a training set, a validation set, and a test set. The division ratio of the dataset is 70% for the training set, 15% for the validation set, and 15% for the test set; Step 5.3.2 The Bayesian optimization algorithm is used to optimize the network model. By constructing a probability model of hyperparameters and selecting the next hyperparameter combination to be evaluated based on the output results of the model, the optimal solution is gradually approximated. Step 5.3.3 The step decay, exponential decay, and adaptive learning rate optimizer are comprehensively used to automatically adjust the learning rate at different stages of training. Step 5.3.4 The early stopping strategy is adopted. When the performance of the validation set does not improve for 10 consecutive epochs, the model training stops, and the model parameters with the best performance on the validation set are retained. Step 5.4 Evaluate and verify the recognition accuracy, generalization ability, and running efficiency of the model: Step 5.4.1 In terms of the recognition accuracy of the model, the imputation accuracy of the model is quantified by three error metrics: root mean square error, mean absolute error, and coefficient of determination. Step 5.4.2 Test the model under different missing rates, including the above three error metrics under data missing rates such as 5%, 10%, 20%, 30%, etc. Step 5.4.3 Use the water service system datasets from different cities, seasons, and climate conditions to evaluate the generalization ability of the model. Step 5.4.4 Use the imputation time of the model, the running efficiency on a single standard server (with 32GB of memory and CPU usage not exceeding 50%), the GPU acceleration efficiency, and the model response time to verify the running efficiency of the model; Step 5.5 Result optimization: After being processed by the neural network model, the system will output multiple results. First, the model marks the detected abnormal data. Second, the model cleans and imputes the missing data, filling the data gaps through prediction. Finally, the model provides a prediction of future water volume.