Water meter reading identification method based on Transform network

Through the water meter reading recognition method based on the Transformer network, a multi-module model of image encoding, transition bit recognition and overall reading recognition is constructed, which solves the problems of low accuracy and great environmental impact of water meter reading recognition in the prior art, and achieves high accuracy and robust automated recognition effects.

CN120220155APending Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510260278.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of water meter readings that have low recognition accuracy, are greatly affected by environmental noise, stains, water mist, light, and the inability to effectively deal with transition positions and pointers.

Method used

The water meter reading recognition method based on the Transformer network is adopted, and the water meter reading model including the image encoding module, the transition bit recognition module and the overall reading recognition module is constructed through image acquisition, labeling, model construction and training. The model uses Transformer Encoder to extract image features and uses Transformer Decoder to identify transition bits and overall readings.

Benefits of technology

It realizes automated water meter reading recognition, improves recognition accuracy and robustness, can effectively handle noise and lighting changes in complex environments, and improves the ability to identify transition bits and pointers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220155A_ABST
    Figure CN120220155A_ABST
Patent Text Reader

Abstract

The invention discloses a Transform network-based water meter reading identification method, which comprises the following steps of: acquiring a water meter image which comprises a character wheel area and a pointer reading area; marking a part of the water meter image, wherein the marking comprises the position and reading of a water meter character wheel and the position and reading of a pointer; constructing a water meter reading model, wherein the water meter reading model comprises an image coding module, a transition bit recognition module and an overall reading recognition module; training the water meter reading model by adopting the annotated image and the unannotated image in the image data set to obtain an optimal water meter reading model; and inputting the acquired water meter image and the highest pointer reading into the optimal water meter reading model to obtain the water meter reading. According to the method, a large number of pictures are collected from a real scene, part of data is labeled in detail, and the robustness and generalization of the model are improved through data enhancement and self-training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method for identifying water meter readings based on a Transformer network. Background Art

[0002] A water meter is an instrument for recording the usage of water resources, which is widely used in daily life, industrial and agricultural production. The annual output of various water meters in China exceeds 100 million. In the early days, mechanical water meters occupied the mainstream market due to their reliability, durability, low price, etc. With the passage of time, the stock of mechanical water meters is huge. The dial of a mechanical water meter includes a digit wheel area and a pointer area, and the numbers need to be read and recorded manually. Manual meter reading is time-consuming and laborious, and there are often problems such as misreading and missing reading, and there are also a series of safety hazards and management problems, which need to be upgraded and transformed.

[0003] Some units upgrade by directly replacing with new digital meters, but this method has some problems. For example, replacing the meter requires disassembling the pipeline, which has a wide range of impacts, is difficult to plan, and brings inconvenience to users; the new digital meters are expensive, difficult to maintain, and the cost has increased significantly.

[0004] The method based on computer vision is a solution widely used in the industry at present. It uses a camera to collect meter images and transmit them, and then uses computer vision technology to obtain the meter readings. This solution has the advantages of low cost, ready-to-use, easy to collect and manage, good scalability, etc., but there are still some problems to be solved.

[0005] The installation environment of water meters is complex and diverse, with serious noise, stains, water mist, light, etc. affecting recognition. There are transition positions when the digit wheels of water meters rotate. Some water meters are loose due to poor gear engagement and cannot be recognized simply as ordinary characters. In addition, due to the presence of pointers, the overall reading of the digit wheels needs to be determined in combination with the highest digit of the pointer. Summary of the Invention

[0006] In order to overcome the above-mentioned disadvantages and deficiencies of the prior art, the purpose of the present invention is to provide a method for identifying water meter readings based on a Transformer network.

[0007] The purpose of the present invention is achieved by the following technical solutions:

[0008] A method for identifying water meter readings based on a Transformer network includes:

[0009] Collecting a water meter image, where the water meter image includes a digit wheel area and a pointer reading area;

[0010] Labeling some of the water meter images, where the labeling includes the positions and readings of the water meter digit wheels and the positions and readings of the pointers;

[0011] Build a water meter reading model, where the water meter reading model includes an image encoding module, a transition digit recognition module, and an overall reading recognition module;

[0012] Use the labeled images and unlabeled images in the image dataset to train the water meter reading model to obtain an optimal water meter reading model;

[0013] Input the collected water meter image and the highest digit pointer reading into the optimal water meter reading model to obtain the water meter reading.

[0014] Furthermore, in the water meter image, the water meter dial part is centered and facing the acquisition lens directly.

[0015] Furthermore, label the positions of the water meter wheels and perform cropping. After cropping, the water meter wheel pictures are scaled, grayscaled, and then data-augmented.

[0016] Furthermore, the image encoding module uses a Transformer Encoder to extract image features and encode the image information.

[0017] Furthermore, the picture is preprocessed at the input layer, and then the vector dimensions are transformed through an embedding layer and a flattening layer. Position information is added at the position encoding layer, and finally, through six Transformer Encoder Layers, the image features are obtained.

[0018] Furthermore, for the transition digit recognition module: the transition digit sequence is converted into a vector through an embedding layer, position information is added at the position encoding layer, and then cross-attention is calculated with the image features through three Transformer Decoder Layers. Finally, the transition digit classification probability is obtained through a classification layer.

[0019] Furthermore, for the overall reading recognition module: the transition digit classification probability and the highest digit pointer reading are concatenated as the input, converted into a vector through an embedding layer, position information is added at the position encoding layer, cross-attention is calculated with the image features through three Transformer Decoder Layers, and finally, the classification probability is obtained through a classification layer. When the overall reading recognition module runs, the input is masked by bit and decoded in parallel to obtain the final overall reading.

[0020] Furthermore, use the labeled images and unlabeled images in the image dataset to train the water meter reading model. The training process of the transition digit recognition model is as follows:

[0021] Train with images with real labels to obtain a teacher model;

[0022] Input the unlabeled images into the teacher model, predict to obtain pseudo-labels, and select the pseudo-labels with high confidence;

[0023] The pseudo-label data obtained by screening is combined with the real label data to train a student model;

[0024] Taking the student model as the teacher model, iterate the steps of predicting pseudo-labels and training the student model. When the number of iterations reaches the preset value, a transition position recognition model is obtained.

[0025] Furthermore, when training the overall reading recognition model, fix the weights of the transition position recognition model and train for a preset number of times to obtain the weights of the overall reading recognition model.

[0026] Furthermore, use the labeled images and unlabeled images in the image dataset to train the water meter reading model to obtain the optimal water meter reading model;

[0027] During the training process, cross-entropy is used as the loss function, calculate the loss bit by bit for the readings, and perform backpropagation layer by layer, so as to update the model parameters to train the optimal water meter reading model;

[0028] During training, the AdamW method is used as the optimizer.

[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0030] (1) Compared with traditional manual meter reading, the present invention can realize automatic instrument reading recognition, save a large amount of labor costs, and has obvious advantages in aspects such as batch data processing;

[0031] (2) The present invention adopts an encoder-decoder model structure based on Transformer, which can better extract image features and fuse context information, and has a higher degree of parallelism. Explicitly model the transition position to obtain the transition position reading. By introducing more strict supervision information, the accuracy of the dial reading is improved. Add the highest digit pointer information, combine the transition position context information and image features to obtain the accurate overall reading. When decoding the overall reading, perform bit masking and parallel decoding, thereby improving the inference speed.

[0032] (3) The present invention collects a large number of pictures from real scenes, and makes detailed annotations for some data. Through data augmentation and self-training, the robustness and generalization of the model are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is the overall flow block diagram of the present invention.

[0034] Figure 2 It is the structural schematic diagram of the network model of the present invention.

[0035] Figure 3 It is the schematic diagram of the transition position recognition model of the present invention.

[0036] Figure 4 This is a schematic diagram of the overall reading recognition model of the present invention. Specific embodiments

[0037] The following combines embodiments to further elaborate on the present invention in detail, but the implementation manners of the present invention are not limited thereto.

[0038] Embodiment

[0039] As Figure 1 shown, a water meter reading recognition method based on the Transformer network includes processes such as image acquisition, data processing, model construction, model training, and reading recognition.

[0040] The specific steps are as follows:

[0041] S1 Image acquisition, specifically using a camera device such as a mobile phone or dedicated hardware to capture an image of the water meter dial.

[0042] In a real application scenario, capture an image of the water meter dial. When shooting, it should include the digit wheel and pointer reading areas, and try to ensure that the lens is perpendicular to the dial to reduce parallax. The collected pictures should cover all specifications of water meters included in the application, and the total number of images should be no less than 150,000.

[0043] S2 Data processing: Label some of the collected images. The labeled content is the positions and readings of the digit wheels and pointers of the water meter. Among them, the digit wheel readings should include the transitional position readings and the overall readings, and crop the images according to the positions.

[0044] Furthermore, the ratio of labeled to unlabeled is 1:10; in this embodiment, there are 15,000 labeled and 150,000 unlabeled.

[0045] Among them, for the transitional position of the digit wheel, the labeling format is Xv. For example, if a certain digit wheel is between 0 and 1 at the transitional position, it is labeled as "0v". There are a total of 20 types of labels for the 10 digits from 0 to 9 plus 10 transitional positions. For the overall pointer reading, the specific reading of each digit wheel is obtained according to the actual carry rule, and there are 10 types from 0 to 9 for each digit. For the pointer reading, only the reading of the highest digit pointer needs to be labeled, and there are 10 types from 0 to 9 in total.

[0046] Use a quadrilateral to label the position of the digit wheel, record the four vertex coordinates of the quadrilateral, crop the digit wheel image after labeling, and scale the digit wheel image to a unified size.

[0047] Expand the image samples by performing random data augmentation on the digit wheel images. Among them, the random augmentation includes random rotation, random affine transformation, perspective transformation, brightness transformation, Gaussian noise, dynamic blur, etc. After augmentation, perform channel normalization on the images.

[0048] Construction of the S3 water meter reading model is as follows Figure 2 As shown in the figure, the water meter reading model includes an image encoding module, a transition digit recognition module, and an overall reading recognition module. Among them, the image encoding module extracts image information through Vision Transformer and encodes the image information; the transition digit recognition module autoregressively obtains the transition digit reading according to the image features through the Transformer decoder; the overall reading recognition module combines the above-mentioned image features, transition digit reading, and the highest digit pointer reading to obtain the final overall reading.

[0049] (1) The specific structure of image encoding is shown in Table 1

[0050] Table 1

[0051]

[0052]

[0053] Image encoding module: The picture is preprocessed in the input layer, and then the vector dimension is transformed through the embedding layer and the flattening layer. Position information is added in the position encoding layer, and finally, through six layers of Transformer Encoder Layer, image features of size 256×192 are obtained.

[0054] As Figure 3 shown, x ∈ R 32×128×1 The picture of size is divided into x p ∈ R 256×4×4 image patches, where (32, 128, 1) are the height, width, and number of channels of the picture respectively, 4×4 is the size of the image patch, and N = (32×128) / (4×4) = 256. The image is flattened after convolution with 4×4 to a vector of size 256×192. 192 is the size of the Transformer dimension. After passing through the embedding layer, it is added to the learnable position encoding and sent into the Transformer Encoder to extract and encode the image information, obtaining f v ∈ R 256×192 .

[0055] (2) The structure of the transition digit recognition module is shown in Table 2

[0056] Table 2

[0057]

[0058]

[0059] Transition position recognition module: The transition position sequence is converted into a vector through the embedding layer, and position information is added in the position encoding layer. Then, cross-attention is calculated with the image features through three layers of Transformer Decoder Layer. Finally, the transition position classification probability is obtained through the classification layer. This process is repeated multiple times, and the classification probabilities of each position are obtained autoregressively.

[0060] As Figure 3 shown, the transition position recognition module uses Transformer Decoder for autoregressive decoding, and the transition position classification probability p t ∈R 9×23 is obtained through the classification layer and the Softnax layer, where 9 is the maximum decoding length and 23 is the number of transition position categories. It includes 20 types of transition positions and BOS, EOS, Padding.

[0061] During training and prediction, masking is used to ensure that only the previous results can be seen at each time step, and future results cannot be seen, so as to ensure that the model has the same input information during training and prediction and avoid information leakage during training.

[0062] (3) The structure of the overall reading recognition module is shown in Table 3:

[0063] Table 3

[0064]

[0065]

[0066] Overall reading recognition module: The transition position classification probability is concatenated with the highest digit reading of the pointer as the input, which is converted into a vector through the embedding layer, and position information is added in the position encoding layer. Cross-attention is calculated with the image features through three layers of Transformer Decoder Layer. Finally, the classification probability is obtained through the classification layer. When this module runs, the input is bit-masked and decoded in parallel to obtain the final overall reading.

[0067] As Figure 4 shown, the overall reading recognition model combines the image features, the transition position context information, and the highest digit reading of the pointer to obtain the overall reading of the dial. The highest digit reading of the pointer is one-hot encoded to obtain p p ∈R 1×23 , which is concatenated with the transition position classification probability as the input p∈R 10×23 . After passing through the embedding layer, the input is added with position encoding to obtain t∈R 10×192 :

[0068] p = Concat(p p , p t )

[0069] t = pE + E pos

[0070] Use the Transformer Decoder to decode the overall reading, and obtain the text feature f from t through the multi-head self-attention layer t , at this time, t is used as Query, Key, and Value. Then, calculate the cross-attention between the text feature and the image feature, and the text feature f t is used as Query, and the image feature f v is used as Key and Value, and the obtained result passes through the multi-layer perceptron and the linear layer to obtain the reading classification probability y ∈ R 10×13 , 10 is the decoding length, and 13 is the number of overall reading categories, including 0-9 and BOS, EOS, Padding:

[0071]

[0072] y = Linear(MLP(f) + f)

[0073] When decoding the overall reading, copy the text feature t M times to obtain the feature matrix T ∈ R 10×10×192 , and perform bit masking, that is, mask according to the diagonal of the first and second dimensions of the matrix, so as to achieve parallel decoding. As shown in Table 4, p represents the pointer information, and t n represents the nth bit of pointer information, 0 represents being covered by the mask, and 1 represents the available information. This means that when decoding each bit result, it can utilize the transitional bit context information and pointer information except itself, and combine with the image feature to obtain the accurate final reading.

[0074] Table 4

[0075]

[0076] S4. Model training: Use the labeled images and unlabeled images in the image dataset to train the water meter reading model to obtain the optimal water meter reading model.

[0077] 1. Initialization: Kaiming initialization.

[0078] 2. When training the transitional bit recognition model, use the self-training method:

[0079] (1): Use the labeled images to train the teacher model;

[0080] (2): Input the unlabeled images into the teacher model, predict the pseudo-labels, and filter according to the confidence;

[0081] (3): Combine the filtered pseudo-label data with the labeled data to train the student model;

[0082] (4) Use the student model as the teacher model to iterate steps (2) and (3).

[0083] Iterate the above steps three times, with 30 rounds of training each time, to obtain the transition position recognition model.

[0084] When training the overall reading recognition model, fix the weights of the transition position recognition model and train for 50 rounds to obtain the weights of the overall reading recognition model.

[0085] 3. Training parameters:

[0086] Number of iterations: 30;

[0087] Optimizer: Use the AdamW method;

[0088] Learning rate: 0.0005;

[0089] Learning rate update strategy: Use the cosine function (CosineAnnealingLR);

[0090] Weight decay coefficient: 0.01.

[0091] 4. The cross-entropy loss function is used in the training process, and the model parameters are updated through the gradient backpropagation algorithm to realize the training of the network model.

[0092] S5. Reading recognition: Input the collected water meter image and the highest digit pointer reading into the optimal water meter reading model to obtain the water meter reading.

[0093] The instrument reading recognition process specifically includes: (1) Input an instrument reading image into the network, use the trained recognition model and parameters to predict and recognize the logarithm to obtain the recognition result; (2) Return the recognition result and calculate the arithmetic mean of the recognition confidence of each digit reading to obtain the recognition confidence.

[0094] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A water meter reading recognition method based on Transformer network, characterized in that: include: Collecting a water meter image, wherein the water meter image includes a character wheel area and a pointer reading area; Annotating a portion of the water meter image, wherein the annotation includes the water meter wheel position and reading and the pointer position and reading; Constructing a water meter reading model, the water meter reading model includes an image encoding module, a transition position recognition module and an overall reading recognition module; The water meter reading model is trained using labeled images and unlabeled images in the image dataset to obtain the optimal water meter reading model. The collected water meter image and the highest digit pointer reading are input into the optimal water meter reading model to obtain the water meter reading.

2. The water meter reading recognition method according to claim 1, characterized in that: In the water meter image, the water meter dial is centered and faces the acquisition lens.

3. The water meter reading recognition method according to claim 1, characterized in that: The position of the water meter wheel is marked and cropped. After cropping, the water meter wheel image is scaled and grayscaled for data enhancement.

4. The water meter reading recognition method according to claim 1, characterized in that: The image encoding module uses Transformer Encoder to extract image features and encode image information.

5. The water meter reading recognition method according to claim 4, characterized in that: The image is preprocessed in the input layer, and then the vector dimension is transformed through the embedding layer and the flattening layer. The position information is added in the position encoding layer, and finally the image features are obtained through the six-layer Transformer Encoder Layer.

6. The water meter reading recognition method according to claim 1, characterized in that: The transition bit recognition module: the transition bit sequence is converted into a vector through the embedding layer, the position information is added in the position encoding layer, and then the cross attention is calculated with the image features through the three-layer TransformerDecoder Layer, and finally the transition bit classification probability is obtained through the classification layer.

7. The water meter reading recognition method according to claim 1, characterized in that: The overall reading recognition module: the transition bit classification probability and the pointer highest bit reading are spliced ​​as input, converted into a vector through the embedding layer, the position information is added in the position encoding layer, the cross attention is calculated through the three-layer Transformer Decoder Layer and the image features, and finally the classification probability is obtained through the classification layer. When the overall reading recognition module is running, the input is bit-masked and decoded in parallel to obtain the final overall reading.

8. The water meter reading recognition method according to claim 1, characterized in that: The water meter reading model is trained using the annotated images and unannotated images in the image dataset. The training process of the transition position recognition model is as follows: The teacher model is trained using images with real labels; Input the unlabeled image into the teacher model, predict the pseudo label, and filter out the high-confidence pseudo label; The pseudo-label data obtained by screening is combined with the real label data to train the student model; The student model is used as the teacher model, and the pseudo-label prediction step and the training step are iterated. When the number of iterations reaches a preset value, a transition position recognition model is obtained.

9. The water meter reading recognition method according to claim 1, characterized in that: When training the overall reading recognition model, the transition position recognition model weight is fixed, and the overall reading recognition model weight is obtained by training for a preset number of times.

10. The water meter reading recognition method according to any one of claims 1 to 9, characterized in that: The water meter reading model is trained using labeled images and unlabeled images in the image dataset to obtain the optimal water meter reading model. The training process uses cross entropy as the loss function, calculates the loss of the reading bit by bit, and transfers it back layer by layer to update the model parameters to train the optimal water meter reading model; The optimizer uses the AdamW method during training.