Ammeter reading error correction method based on deep learning
Through deep learning's rotation-invariant position encoding and three-dimensional cross-attention mechanism, combined with confidence-driven bimodal enhanced reasoning, the problems of recognition accuracy and robustness of the meter reading system in complex environments are solved, and efficient meter reading error correction is achieved.
Patent Information
- Application Number
- CN202510911028.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-21
AI Technical Summary
The existing electricity meter reading system is prone to digital position mismatch and stroke confusion in complex environments, and old meters are easily affected by glass reflections, pointer occlusions and background noise, resulting in insufficient recognition accuracy and robustness.
A deep learning-based electricity meter reading error correction method is adopted, which utilizes rotation-invariant position encoding and three-dimensional cross-attention mechanism, combined with confidence-driven bimodal enhanced reasoning to achieve highly robust recognition of electricity meter images.
It significantly improves the accuracy of digital recognition in multi-angle shooting scenarios of electricity meters, reduces the risk of missed and misjudgment, and improves the robustness and interpretability of readings in noisy environments.
Smart Images

Figure CN120823587A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric meters, and in particular to an electric meter reading error correction method based on deep learning. Background Art
[0002] With the rapid development of smart grids and the power Internet of Things (IoT), automated meter reading and intelligent inspection have become crucial components of energy metering. Existing meter reading systems primarily rely on image recognition algorithms based on two-dimensional convolutional neural networks or traditional optical character recognition to process meter plate images. In practice, a large number of meters are installed in shafts, closets, and other high-rise, complex environments. Manual or robotic image acquisition often results in image distortion from tilt, rotation, and mirroring due to limited space and operational difficulties. Traditional recognition methods based on fixed grids or simple relative position encoding are prone to digit mismatches and stroke confusion when dealing with rotation and meter viewing angle changes. This is particularly true for easily confused digits such as "1 / 7" and "5 / 6." Furthermore, older or composite meters often suffer from interference from glass reflections, pointer obstructions, and background dust. Two-dimensional feature modeling is limited in its ability to decouple digits, pointers, and background noise, leading to systematic misjudgments and failing to achieve a balanced recognition accuracy and robustness. Summary of the Invention
[0003] One object of the present invention is to propose a method for correcting meter readings based on deep learning. The present invention ensures that highly robust integrated reasoning can be automatically triggered in low-confidence and abnormal change scenarios, significantly reducing the risks of missed judgments and misjudgments.
[0004] A method for correcting electric meter reading errors based on deep learning according to an embodiment of the present invention includes the following steps:
[0005] Acquire target electric meter image data and preprocess it to obtain standardized electric meter image data and a first preprocessing parameter set;
[0006] performing a rough estimation of the rotation angle of the standardized electricity meter image data according to the first preprocessing parameter set to generate a rotation angle estimation parameter, dividing the standardized electricity meter image data into a sequence of image blocks having overlapping areas, and generating a rotation-invariant position vector for each image block to form a sequence of image blocks having the rotation-invariant position vector;
[0007] The image block sequence with rotation-invariant position vectors is input into the improved Transformer encoder, the encoding feature representation is extracted using the self-attention mechanism, and the first encoding feature representation is output;
[0008] Insert a three-dimensional cross-attention module between the improved Transformer encoder output and the Transformer decoder input to output a three-dimensional cross-attention feature sequence;
[0009] The three-dimensional cross-attention feature sequence is input into the Transformer decoder, and the preliminary meter digital reading result is generated by combining the historical meter reading prior information. The confidence evaluation is performed on the preliminary meter digital reading result to obtain the confidence evaluation result;
[0010] When the confidence assessment result is lower than the preset threshold or the difference with the prior information of historical meter readings is greater than the preset deviation range, the dual-modal enhanced reasoning mechanism is triggered, and re-reasoning is performed based on the standardized meter image data and its mirror image to output the enhanced meter digital reading result;
[0011] A result selection is performed between the enhanced electric meter digital reading result and the preliminary electric meter digital reading result based on the confidence evaluation result to obtain a final electric meter digital reading result.
[0012] Optionally, the standardized electricity meter image data and the first preprocessing parameter set are constructed, including:
[0013] Collect target electric meter image data;
[0014] The brightness deviation is obtained by subtracting the overall average brightness value of the target meter image data from the original pixel intensity value of each pixel point of the target meter image data, and the difference between the original image pixel value and the brightness deviation is calculated to obtain the target meter image after illumination balance;
[0015] The target meter image I after illumination balance eq Perform reflection interference suppression and construct the cosine similarity function ρ based on the image gradient vector and the pose vector r (x, y), extract the local high-reflection area mask and perform reflection weight suppression filtering to obtain the target meter image after reflection suppression;
[0016] A perspective transformation matrix is constructed based on the Euler rotation angle components in the collected pose information. The target meter image after reflection suppression is mapped from the original perspective to an approximately orthographic perspective image. The perspective transformation obtains standardized meter image data by transforming the spatial coordinates of the original image with the perspective transformation matrix.
[0017] The standardized electric meter image data and the first preprocessing parameter set are output.
[0018] Optionally, generating a rotation-invariant position vector for each image block includes:
[0019] Perform directional statistics on the global gradient field of the standardized electric meter image data according to the perspective transformation matrix in the first preprocessing parameter set and the consistency measurement result of the reflection direction and the observation direction, and calculate the overall main direction angle of the target electric meter image;
[0020] Dividing the standardized electric meter image data into a plurality of image block sequences with overlapping areas in a two-dimensional space;
[0021] The center position of each image block is converted into polar coordinates based on the image center point as the reference origin;
[0022] Based on the polar coordinate expression, the distance component and the angle component of each image block are further mapped into a complex position vector to obtain a rotation-invariant position vector;
[0023] Each image block P k The corresponding rotation invariant position vector RIPE k Splicing to form a sequence of image blocks with rotation-invariant position encoding.
[0024] Optionally, the improved Transformer encoder includes:
[0025] The image block sequence with rotation-invariant position encoding is input into the improved Transformer encoder. The visual feature vector of each image block with rotation-invariant position encoding is concatenated with its rotation-invariant position vector. The local adaptive query vector is calculated through linear mapping and layer normalization.
[0026] The visual similarity is calculated by comparing the visual feature vector with the visual feature vectors of all adjacent image blocks to generate the neighborhood mask weight.
[0027] After concatenating the visual feature vector of each image block with the rotation-invariant position vector, a key vector and a value vector of the structural constraint are generated through a linear mapping function.
[0028] Using locally adaptive query vectors Key vector and value vector with structural constraints Computational Improvement of Transformer Encoder Structure to Enhance Self-Attention Weights
[0029]
[0030] Where d is the dimension of the key vector, To improve the structure of the Transformer encoder, the self-attention weight is enhanced to indicate the degree to which the k-th image block pays attention to the structural information of the j-th image block during the encoding process;
[0031] The structure-enhanced self-attention weights are used to perform weighted summation on the value vectors of the structural constraints of all image blocks to obtain the first encoded feature representation of the structure enhancement.
[0032] Optionally, the output three-dimensional cross-attention feature sequence includes:
[0033] Reconstructing the structure-enhanced first encoded feature representation set into a three-dimensional feature tensor;
[0034] In the spatial dimension, channel dimension, and depth or time dimension of the three-dimensional feature tensor, the query vector, key vector, and value vector of the spatial dimension, channel dimension, and depth or time dimension are calculated for each feature block respectively;
[0035] Calculate the spatial attention weight A based on the query vector and key vector in the spatial dimension, channel dimension, and depth or time dimension s , channel attention weight A c and the depth or temporal attention weight A d ;
[0036] Perform weighted aggregation on the attention weights to form a three-dimensional cross-attention output feature tensor F 3D :
[0037] F 3D =α s ·(A s ·V s )+α c ·(A c ·V c )+α d ·(A d ·V d );
[0038] Among them, α s , α c , α d is the learnable weight factor in three dimensions: space, channel and depth, V s 、V c 、V d are value vectors of spatial dimension, channel dimension, and depth or time dimension respectively;
[0039] The 3D cross-attention output feature tensor is flattened into a 3D cross-attention feature sequence according to the spatial arrangement order of the image blocks.
[0040] Optionally, performing confidence assessment on the preliminary electric meter digital reading result includes:
[0041] For each feature vector in the three-dimensional cross-attention feature sequence, a set of decoder query vectors is calculated using linear transformation and layer normalization.
[0042] The historical meter reading prior information sequence is transformed into a prior embedding sequence through an embedding matrix and sequentially concatenated with the decoder query vector set to form a joint decoding input sequence. The joint decoding input sequence is used to simultaneously introduce current features and historical prior information during the decoding process.
[0043] In each layer of the Transformer decoder, the joint decoding input sequence is transformed into the value vector required for self-attention through the projection matrix, and the value vector is weighted and summed to obtain the output features of the decoding layer;
[0044] Based on the output features of the last layer of decoder, a position-sensitive digital classification head is connected to generate a digital category logarithmic vector z for each sequence position. t =[z t,0 ,z t,1 ,…,z t,9 ], calculate the digital probability distribution by adjusting the Softmax function by temperature:
[0045]
[0046] Among them, T>0 is an adjustable temperature coefficient used to control the smoothness of the probability distribution;
[0047] For each sequence position, the digital category with the maximum probability in the digital probability distribution is selected as the preliminary meter digital reading result, and the maximum probability is used as the confidence value of the current position;
[0048] All confidence values in the confidence vector are averaged to obtain an overall confidence evaluation result. When the overall confidence evaluation result is greater than or equal to a preset confidence threshold, a high confidence evaluation result is output; otherwise, a low confidence evaluation result is output.
[0049] Optionally, the conditions for triggering the bimodal enhanced reasoning mechanism include:
[0050] When the overall confidence assessment result is less than the preset confidence threshold or the difference between the preliminary meter digital reading result and the historical meter reading prior information sequence is greater than the preset deviation range threshold, the dual-modal enhanced reasoning mechanism is automatically triggered.
[0051] Optionally, the bimodal enhanced reasoning mechanism includes:
[0052] For the standardized meter image data, the original perspective input and the corresponding mirror perspective input are constructed respectively;
[0053] Calculate the original perspective enhanced reading results and the mirror perspective enhanced reading results of the original perspective input and the mirror perspective input, and calculate the corresponding confidence evaluation results respectively;
[0054] Compare the confidence assessment results of the corresponding positions of the original perspective enhanced reading result and the mirror perspective enhanced reading result. If the confidence assessment result of the original perspective enhanced reading result is not less than the confidence assessment result of the mirror perspective enhanced reading result, the final enhanced meter digital reading result uses the original perspective enhanced reading result; otherwise, the mirror perspective enhanced reading result is used. The final enhanced meter digital reading result is the set of numbers at each position obtained by confidence comparison.
[0055] Optionally, the calculation of the final meter digital reading result includes:
[0056] Calculating the confidence value of each digit of the preliminary meter digital reading result and the enhanced meter digital reading result respectively;
[0057] For each bit of the digital sequence, a number with a confidence value greater than a threshold is selected from the preliminary meter digital reading result and the enhanced meter digital reading result as the final meter digital reading result of that bit. The final meter digital reading result is composed of the predicted numbers with a confidence value greater than the threshold at each position.
[0058] The beneficial effects of the present invention are:
[0059] (1) The position coding of the present invention breaks through the limitation of rotation disturbance and greatly improves the accuracy of digital recognition in the multi-angle shooting scenario of the electric meter. The spatial coordinates of each image block in the electric meter image are converted into polar coordinate expression by using rotation-invariant position coding, and the spatial position is rotated and aligned by complex vectors, eliminating the sensitivity of absolute or relative position coding in conventional Transformer to changes in shooting angle. Regardless of the tilt, rotation or mirroring when the electric meter is photographed, the model can map the spatial distribution of the same number in different postures into a unified feature representation, significantly reducing the problems of digital sequence mismatch and confusion of similar strokes.
[0060] (2) The three-dimensional cross-attention mechanism of the present invention realizes the fusion of space-channel-depth multimodal information, finely decouples complex noise, and effectively suppresses system misreadings such as pointer occlusion and glass reflection. A three-dimensional cross-attention module is introduced between the Transformer encoder and decoder to construct independent query, key and value vectors in space, channel and depth or time dimensions respectively. The three-dimensional weighted fusion fully captures the fine-grained structural information of the meter image in different modalities, significantly improving the model's ability to distinguish multi-source interference and improving the robustness and interpretability of readings in noisy environments.
[0061] (3) The present invention designs a confidence-driven dual-modal enhanced reasoning and dynamic decision-making mechanism to comprehensively improve the credibility of error correction results and realize the self-correction capability of the model. After each reasoning, the confidence corresponding to the preliminary digital sequence and the enhanced digital sequence (original perspective and mirror perspective) is automatically calculated, and bit-by-bit optimal output is achieved based on dynamic standards, ensuring that highly robust integrated reasoning can be automatically triggered in low-confidence and abnormal change scenarios, significantly reducing the risk of missed judgments and misjudgments. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0063] Figure 1 This is a flowchart of a meter reading error correction method based on deep learning proposed by the present invention. DETAILED DESCRIPTION
[0064] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0065] refer to Figure 1 , a meter reading error correction method based on deep learning, comprising the following steps:
[0066] Acquire target electric meter image data and preprocess it to obtain standardized electric meter image data and a first preprocessing parameter set;
[0067] In this embodiment, the standardized electricity meter image data and the first preprocessing parameter set are constructed, including:
[0068] Collect target electric meter image data;
[0069] The target meter image data is collected by an image acquisition module with multimodal acquisition function. The image acquisition module obtains the shooting angle, position and posture information corresponding to the target meter image data to form the collection posture information. The collection posture information includes the position component of the collection position in the three-axis spatial coordinates and the Euler rotation angle components around the X, Y and Z axes.
[0070] The brightness deviation is obtained by subtracting the overall average brightness value of the target meter image data from the original pixel intensity value of each pixel point of the target meter image data, and the difference between the original image pixel value and the brightness deviation is calculated to obtain the target meter image after illumination balance;
[0071] Brightness deviation is used to measure the difference between each pixel and the overall average brightness of the target meter image data.
[0072] The target meter image I after illumination balance eq Perform reflection interference suppression and construct the cosine similarity function ρ based on the image gradient vector and the pose vector r (x, y), extract the local high-reflection area mask, and perform reflection weight suppression filtering to obtain the target meter image after reflection suppression:
[0073]
[0074] in, is the gradient vector of the target meter image at pixel (x, y), is the pose vector calculated based on the collected pose information, ρ r (x, y) is the cosine similarity function, which is used to measure the consistency between the reflection direction and the observation direction;
[0075] A perspective transformation matrix is constructed based on the Euler rotation angle components in the collected pose information. The target meter image after reflection suppression is mapped from the original perspective to an approximately orthographic perspective image. The perspective transformation obtains standardized meter image data by transforming the spatial coordinates of the original image with the perspective transformation matrix.
[0076] Output standardized electricity meter image data and a first preprocessing parameter set, where the first preprocessing parameter set includes an overall average brightness value of the target electricity meter image data, a mask of a local high-reflection area, a measurement result of consistency between the reflection direction and the observation direction, and a perspective transformation matrix.
[0077] performing a rough estimation of the rotation angle of the standardized electricity meter image data according to the first preprocessing parameter set to generate a rotation angle estimation parameter, dividing the standardized electricity meter image data into a sequence of image blocks having overlapping areas, and generating a rotation-invariant position vector for each image block to form a sequence of image blocks having the rotation-invariant position vector;
[0078] In this embodiment, generating a rotation-invariant position vector for each image block includes:
[0079] Perform directional statistics on the global gradient field of the standardized electric meter image data according to the perspective transformation matrix in the first preprocessing parameter set and the consistency measurement result of the reflection direction and the observation direction, and calculate the overall main direction angle of the target electric meter image;
[0080] The overall main direction angle is used to represent the angle between the meter digit arrangement direction and the horizontal direction.
[0081] Dividing the standardized electric meter image data into a plurality of image block sequences with overlapping areas in a two-dimensional space;
[0082] When dividing, the center position of each image block is determined according to the image size and the sliding step size. There is an overlapping area between every two adjacent image blocks. The overlapping area ensures the continuity of information between different image blocks. The number of image blocks is determined by the overall image size and the set division rules.
[0083]
[0084] Where N is the number of image blocks, O k represents the kth image block P k With the k+1th image block P k+1 The overlapping area between them.
[0085] The center position of each image block is converted into polar coordinates based on the image center point as the reference origin;
[0086] Polar coordinates include the distance component between the center position of each image block and the center point of the image, as well as the angle component between the center position and the center point of the image. The angle component achieves angle alignment by removing the influence of the rotation angle estimation parameters. The distance component and the angle component are used together to characterize the relative position relationship of each image block in space.
[0087] Based on the polar coordinate expression, the distance component and the angle component of each image block are further mapped into a complex position vector to obtain a rotation-invariant position vector;
[0088] The complex position vector contains a real part and an imaginary part, which are used to encode the position features of the image block in the rotation alignment space. The real part represents the horizontal features of the spatial coordinates, and the imaginary part represents the vertical features of the spatial coordinates. The rotation-invariant position vector is a vector composed of the real part and the imaginary part of the complex position vector.
[0089] Each image block P k The corresponding rotation invariant position vector RIPE k Splicing to form a sequence of image blocks with rotation-invariant position encoding
[0090] The position encoding in this embodiment breaks through the limitation of rotational disturbance and greatly improves the accuracy of digital recognition in multi-angle shooting scenarios of electric meters. The spatial coordinates of each image block in the electric meter image are converted into polar coordinate expression using rotation-invariant position encoding, and the spatial position is rotated and aligned through complex vectors, eliminating the sensitivity of absolute or relative position encoding in conventional Transformers to changes in shooting angles. Regardless of the tilt, rotation or mirroring of the electric meter during shooting, the model can map the spatial distribution of the same number in different postures into a unified feature representation, significantly reducing the problems of digital sequence mismatch and confusion of similar strokes.
[0091] The image block sequence with rotation-invariant position vectors is input into the improved Transformer encoder, the encoding feature representation is extracted using the self-attention mechanism, and the first encoding feature representation is output;
[0092] In this embodiment, the Transformer encoder is improved, including:
[0093] The image block sequence with rotation-invariant position encoding is input into the improved Transformer encoder. The visual feature vector of each image block with rotation-invariant position encoding is concatenated with its rotation-invariant position vector. The local adaptive query vector is calculated through linear mapping and layer normalization.
[0094] The locally adaptive query vector is designed based on the visual confusability and spatial distribution irregularity of digital features in the meter reading error correction scenario. The locally adaptive query vector is used to highlight the visual feature differences between different digital regions.
[0095] The visual similarity is calculated by comparing the visual feature vector with the visual feature vectors of all adjacent image blocks to generate the neighborhood mask weight.
[0096] The neighborhood mask weight is used to adaptively select image blocks with similar visual features to the current image block as the neighborhood of the local self-attention mechanism, effectively suppressing the interference of background noise in the meter reading. The visual similarity is compared with the preset threshold. If the visual similarity is greater than or equal to the preset threshold, the neighborhood mask weight is 1, otherwise it is 0.
[0097] After concatenating the visual feature vector of each image block with the rotation-invariant position vector, a key vector and a value vector of the structural constraint are generated through a linear mapping function.
[0098] The key vector and value vector of the structure constraint are used to encode the digital stroke structure information of the image block in the local rotation invariant space.
[0099] Using locally adaptive query vectors Key vector and value vector with structural constraints Computational Improvement of Transformer Encoder Structure to Enhance Self-Attention Weights
[0100]
[0101] Where d is the dimension of the key vector, To improve the structure of the Transformer encoder, the self-attention weight is enhanced to indicate the degree to which the kth image block pays attention to the structural information of the jth image block during the encoding process. This is used to highlight the geometric continuity and consistency of the digit strokes and optimize the distinguishability of the meter digital features.
[0102] The structure-enhanced self-attention weight is used to perform weighted summation on the value vectors of the structural constraints of all image blocks to obtain the first encoding feature representation of the structure enhancement;
[0103] The structure-enhanced first encoded feature representation is used to reflect the visual feature discrimination and structured geometric information for meter reading error correction.
[0104] Insert a three-dimensional cross-attention module between the improved Transformer encoder output and the Transformer decoder input to output a three-dimensional cross-attention feature sequence;
[0105] In this embodiment, the output of the three-dimensional cross attention feature sequence includes:
[0106] Reconstructing the structure-enhanced first encoded feature representation set into a three-dimensional feature tensor;
[0107] The spatial dimension of the three-dimensional feature tensor is determined by the arrangement shape of the image blocks in the horizontal and vertical directions, and the channel dimension is determined by the characteristic length of each encoded feature vector. The three-dimensional feature tensor is used to restore the relative position relationship of the structural enhancement features in the two-dimensional space.
[0108] In the spatial dimension, channel dimension, and depth or time dimension of the three-dimensional feature tensor, the query vector, key vector, and value vector of the spatial dimension, channel dimension, and depth or time dimension are calculated for each feature block respectively;
[0109] The query vector, key vector, and value vector in the spatial dimension are used to extract contextual information between different spatial positions. The query vector, key vector, and value vector in the channel dimension are used to extract feature correlation between different feature channels. The query vector, key vector, and value vector in the depth or time dimension are used to extract information fusion of multi-frame sequences or multimodal inputs in the depth or time dimension. The mapping relationship of each dimension is calculated by different linear mapping weight matrices.
[0110] Calculate the spatial attention weight A based on the query vector and key vector in the spatial dimension, channel dimension, and depth or time dimension s , channel attention weight A c and the depth or temporal attention weight A d ;
[0111] Spatial attention weights are used to measure the feature dependencies between different spatial positions, channel attention weights are used to measure the intrinsic correlation between features of different channels, and depth or temporal attention weights are used to measure the feature synergy of multi-frame or multimodal inputs in the depth or time dimension.
[0112] Perform weighted aggregation on the attention weights to form a three-dimensional cross-attention output feature tensor F 3D The 3D cross-attention output feature tensor integrates multi-dimensional context features at each spatial position, channel, depth, or time point, enabling effective recognition and differentiation of digital boundary structures and background noise in meter image blocks:
[0113] F 3D =α s ·(A s ·V s )+α c ·(A c ·V c )+α d ·(A d ·V d );
[0114] Among them, α s , α c , α d is the learnable weight factor in three dimensions: space, channel and depth, V s 、V c 、V d are value vectors of spatial dimension, channel dimension, and depth or time dimension respectively;
[0115] The 3D cross-attention output feature tensor is flattened into a 3D cross-attention feature sequence according to the spatial arrangement order of the image blocks.
[0116] The three-dimensional cross-attention mechanism in this embodiment realizes the fusion of space-channel-depth multimodal information, finely decouples complex noise, and effectively suppresses system misreadings such as pointer occlusion and glass reflection. A three-dimensional cross-attention module is introduced between the Transformer encoder and decoder to construct independent query, key and value vectors in space, channel and depth or time dimensions respectively. The three-dimensional weighted fusion fully captures the fine-grained structural information of the meter image in different modalities, significantly improving the model's ability to distinguish multi-source interference and improving the robustness and interpretability of readings in noisy environments.
[0117] The three-dimensional cross-attention feature sequence is input into the Transformer decoder, and the preliminary meter digital reading result is generated by combining the historical meter reading prior information. The confidence evaluation is performed on the preliminary meter digital reading result to obtain the confidence evaluation result;
[0118] In this embodiment, a confidence assessment is performed on the preliminary digital meter reading result, including:
[0119] For each feature vector in the three-dimensional cross-attention feature sequence, a set of decoder query vectors is calculated using linear transformation and layer normalization.
[0120] Each decoder query vector is obtained by adding the feature vector to the linear weight matrix and the bias and then normalizing it. The decoder query vector is used to represent the comprehensive contextual features of the meter image block.
[0121] The historical meter reading prior information sequence is transformed into a prior embedding sequence through an embedding matrix and sequentially concatenated with the decoder query vector set to form a joint decoding input sequence. The joint decoding input sequence is used to simultaneously introduce current features and historical prior information during the decoding process.
[0122] In each layer of the Transformer decoder, the joint decoding input sequence is transformed into the value vector required for self-attention through the projection matrix, and the value vector is weighted and summed to obtain the output features of the decoding layer;
[0123] Based on the output features of the last layer of decoder, a position-sensitive digital classification head is connected to generate a digital category logarithmic vector z for each sequence position. t =[z t,0 ,z t,1 ,…,z t,9 ], calculate the digital probability distribution by adjusting the Softmax function by temperature:
[0124]
[0125] Among them, T>0 is an adjustable temperature coefficient used to control the smoothness of the probability distribution;
[0126] For each sequence position, the digital category with the maximum probability in the digital probability distribution is selected as the preliminary meter digital reading result, and the maximum probability is used as the confidence value of the current position;
[0127] The preliminary meter digital reading result is represented as a digital sequence, and the confidence value set is represented as a confidence vector. The confidence vector is used to characterize the credibility of the prediction result of each digit.
[0128] The confidence values in the confidence vector are averaged to obtain an overall confidence evaluation result. When the overall confidence evaluation result is greater than or equal to a preset confidence threshold, a high confidence evaluation result is output; otherwise, a low confidence evaluation result is output.
[0129] The overall confidence evaluation result is used to reflect the confidence level of the entire digital sequence recognition result, and the confidence threshold is used to distinguish between high-confidence and low-confidence recognition results.
[0130] When the confidence assessment result is lower than the preset threshold or the difference with the prior information of historical meter readings is greater than the preset deviation range, the dual-modal enhanced reasoning mechanism is triggered, and re-reasoning is performed based on the standardized meter image data and its mirror image to output the enhanced meter digital reading result;
[0131] In this embodiment, the conditions for triggering the dual-modal enhanced reasoning mechanism include:
[0132] When the overall confidence assessment result is less than the preset confidence threshold or the difference between the preliminary meter digital reading result and the historical meter reading prior information sequence is greater than the preset deviation range threshold, the dual-modal enhanced reasoning mechanism is automatically triggered;
[0133] The difference criterion is obtained by calculating the average value of the absolute difference between the initial meter digital reading and the corresponding position of the historical meter reading prior information sequence. The average value of the absolute difference is used to characterize the overall deviation between the two digital sequences. The average value of the absolute difference is compared with the preset deviation range threshold as the basis for triggering the dual-modal enhanced reasoning mechanism.
[0134] This embodiment is characterized in that the dual-modality enhanced reasoning mechanism includes:
[0135] For the standardized meter image data, the original perspective input and the corresponding mirror perspective input are constructed respectively;
[0136] The mirror view input is obtained by performing a spatial mirror transformation operation on the standardized electricity meter image data. The spatial mirror transformation operation is used to perform a symmetrical transformation on the standardized electricity meter image data along a vertical axis or a horizontal axis, thereby obtaining the mirror view input.
[0137] Calculate the original perspective enhanced reading results and the mirror perspective enhanced reading results of the original perspective input and the mirror perspective input, and calculate the corresponding confidence evaluation results respectively;
[0138] The original perspective enhanced reading result represents the inference output after direct processing of the standardized meter image data, and the mirror perspective enhanced reading result represents the inference output after the mirror perspective input is processed in the same way. The confidence assessment result is used to characterize the credibility level of each inference result.
[0139] Compare the confidence assessment results of the corresponding positions of the original perspective enhanced reading result and the mirror perspective enhanced reading result. If the confidence assessment result of the original perspective enhanced reading result is not less than the confidence assessment result of the mirror perspective enhanced reading result, the final enhanced meter digital reading result uses the original perspective enhanced reading result; otherwise, the mirror perspective enhanced reading result is used. The final enhanced meter digital reading result is the set of numbers at each position obtained by confidence comparison.
[0140] In this implementation, a confidence-driven dual-modal enhanced reasoning and dynamic decision-making mechanism is designed to comprehensively improve the credibility of error correction results and realize the self-correction capability of the model. After each inference, the confidence corresponding to the preliminary digital sequence and the enhanced digital sequence (original perspective and mirror perspective) is automatically calculated, and bit-by-bit optimal output is achieved based on dynamic standards, ensuring that highly robust integrated reasoning can be automatically triggered in low-confidence and abnormal change scenarios, significantly reducing the risk of missed judgments and misjudgments.
[0141] A result selection is performed between the enhanced electric meter digital reading result and the preliminary electric meter digital reading result based on the confidence evaluation result to obtain a final electric meter digital reading result.
[0142] In this embodiment, the calculation of the final meter digital reading result includes:
[0143] The confidence value of each digit of the preliminary meter digital reading result and the enhanced meter digital reading result is calculated respectively; the confidence value of the preliminary meter digital reading result is used to indicate the credibility of each digit of the preliminary prediction, and the confidence value of the enhanced meter digital reading result is used to indicate the credibility of each digit of the enhanced prediction. The confidence value is obtained by comparing the confidence values of the corresponding positions of the confidence vector of the preliminary meter digital reading result and the confidence vector of the enhanced meter digital reading result.
[0144] For each digit in the digital sequence, a digit with a confidence value greater than a threshold is selected from the preliminary meter digital reading result and the enhanced meter digital reading result as the final meter digital reading result of the digit. The final meter digital reading result is used to preferentially select a predicted digit with a confidence value greater than the threshold for each digit. The predicted digit with a confidence value greater than the threshold is preferentially output as the final output.
[0145] Output the final digital meter reading result, which consists of the predicted numbers at each position whose confidence value is greater than the threshold.
[0146] Example 1: In the meter box No. 102 in Building 5 of Community A, visible light + infrared dual-channel images collected by the inspection robot R302 were uploaded to the power data platform. The meter number in the meter box was NJ-B-4301987. The inspection robot was limited in position and space during the collection. The lens was tilted 27° counterclockwise relative to the surface. The sunlight was obliquely incident on the upper left corner of the dial glass. There were reflective spots on the digital area "3" and "8" of the meter, and the lower right side of the number "6" was blocked by the red pointer. The image archive name is "NJ-B-4301987-20240606-091700.jpg".
[0147] Under the traditional CNN-OCR method, the system first preprocesses and recognizes the result, and the output is "38216". The system confidence is marked as 0.64. Data quality inspector Wang Gong manually checked and found that the actual reading on site was "38286". The "8" was mistakenly identified as "1". The reflection + pointer occlusion caused the stroke confusion. Because the confidence was lower than the set threshold (0.75), it was automatically assigned to the manual review queue.
[0148] Subsequently, the same batch of images are intelligently processed using the meter reading error correction method described in the present invention. The images are first illuminated and the reflective area in the upper left corner is automatically suppressed. The system detects the tilt of the shooting angle and automatically calculates the main direction as -27°. All patches are converted into rotationally equidistant space using rotation-invariant position encoding. The patch features are spliced and input into the Transformer encoder. The three-dimensional cross-attention module fuses the infrared depth channel information, suppresses the reflective spots and pointer shadows in the "8" area, and separates the real digital boundaries at the feature layer.
[0149] During the decoding phase, combined with the historical priors of last month's meter reading (last month's reading was "38182"), the preliminary output result is "38286", and the confidence vector is [0.99, 0.96, 0.91, 0.97, 0.93], with an overall confidence of 0.95. All of these are above the system threshold and no manual review is required. The system directly records the data in the general ledger.
[0150] In the image captured by robot R302, the meter number NJ-B-4303159 in Area A on the 7th floor of Community B is installed at a high position, the lens is raised at an angle of 38°, and the number "9" is blocked by dirt. The traditional method outputs "63493" with a confidence level of 0.61. The manual proofreading actually outputs "68493", which is incorrectly judged as "6". After using this embodiment, the system detects high-angle tilt and local dirt, and the rotation-invariant encoding and spatial-depth joint attention significantly separate the digital boundaries. The inference result is "68493" with confidence levels of [0.97, 0.93, 0.90, 0.98, 0.91]. The confidence levels of all positions are higher than the threshold, and the report is correct without review.
[0151] The electric meter number for the underground corridor of Building 7 in Community C is NJ-B-4302346. The inspection robot R303 encountered extreme darkness and fog on the dial during data collection. The traditional method output was "21517" with a confidence level of 0.58. After manual verification, it was found that the actual reading was "21512". The fifth digit "7" should be "2". The error was caused by blurred digital boundaries and background water mist. After adopting the error correction method of the present invention, the system preprocessed the fog image, and rotation-invariant coding + three-dimensional cross-attention accurately identified "2", outputting "21512" with confidence levels of [0.98, 0.93, 0.95, 0.91, 0.95]. The overall confidence level was 0.944, and the result was directly recorded as meeting the standards.
[0152] A random inspection of 3,700 records with a confidence level lower than 0.75 using the traditional method resulted in a final confirmation accuracy of 85.3%. The output confidence levels of the method used in the same batch were all higher than 0.92, and the inspection revealed an accuracy rate of 99.1%. The error-prone scenario data are as follows:
[0153] NJ-B-4301433, severe reflection, traditional method "57177", actual "57171", this embodiment is "57171", confidence level 0.96.
[0154] NJ-B-4301278, the number "1 / 7" is severely worn, the traditional method is "13767", the actual is "13761", and this embodiment is "13761", with a confidence level of 0.95.
[0155] NJ-B-4301980, the pointer blocks the number "6", the traditional method is "43821", the actual is "43861", and this embodiment is "43861", with a confidence level of 0.97.
[0156] NJ-B-4301989, glass stains affect the number "3", the traditional method is "33952", the actual value is "38952", and the value in this embodiment is "38952", with a confidence level of 0.98.
[0157] The method of the present invention only needs to review 1930 groups (review rate 0.82%), which greatly reduces the manual pressure and improves the automation, reliability and economy of the entire inspection cycle.
[0158] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for correcting electric meter reading errors based on deep learning, characterized in that: The steps include: Acquire target electric meter image data and preprocess it to obtain standardized electric meter image data and a first preprocessing parameter set; performing a rough estimation of the rotation angle of the standardized electricity meter image data according to the first preprocessing parameter set to generate a rotation angle estimation parameter, dividing the standardized electricity meter image data into a sequence of image blocks having overlapping areas, and generating a rotation-invariant position vector for each image block to form a sequence of image blocks having the rotation-invariant position vector; The image block sequence with rotation-invariant position vectors is input into the improved Transformer encoder, the encoding feature representation is extracted using the self-attention mechanism, and the first encoding feature representation is output; Insert a three-dimensional cross-attention module between the improved Transformer encoder output and the Transformer decoder input to output a three-dimensional cross-attention feature sequence; The three-dimensional cross-attention feature sequence is input into the Transformer decoder, and the preliminary meter digital reading result is generated by combining the historical meter reading prior information. The confidence evaluation is performed on the preliminary meter digital reading result to obtain the confidence evaluation result; When the confidence assessment result is lower than the preset threshold or the difference with the prior information of historical meter readings is greater than the preset deviation range, the dual-modal enhanced reasoning mechanism is triggered, and re-reasoning is performed based on the standardized meter image data and its mirror image to output the enhanced meter digital reading result; A result selection is performed between the enhanced electric meter digital reading result and the preliminary electric meter digital reading result based on the confidence evaluation result to obtain a final electric meter digital reading result.
2. The electric meter reading error correction method based on deep learning according to claim 1 is characterized in that: The standardized electric meter image data and the first preprocessing parameter set are constructed, including: Collect target electric meter image data; The brightness deviation is obtained by subtracting the overall average brightness value of the target meter image data from the original pixel intensity value of each pixel point of the target meter image data, and the difference between the original image pixel value and the brightness deviation is calculated to obtain the target meter image after illumination balance; The target meter image I after illumination balance eq Perform reflection interference suppression and construct the cosine similarity function ρ based on the image gradient vector and the pose vector r (x, y), extract the local high-reflection area mask and perform reflection weight suppression filtering to obtain the target meter image after reflection suppression; A perspective transformation matrix is constructed based on the Euler rotation angle components in the collected pose information. The target meter image after reflection suppression is mapped from the original perspective to an approximately orthographic perspective image. The perspective transformation obtains standardized meter image data by transforming the spatial coordinates of the original image with the perspective transformation matrix. The standardized electric meter image data and the first preprocessing parameter set are output.
3. The electric meter reading error correction method based on deep learning according to claim 1 is characterized in that: Generating a rotation-invariant position vector for each image block includes: Perform directional statistics on the global gradient field of the standardized electric meter image data according to the perspective transformation matrix in the first preprocessing parameter set and the consistency measurement result of the reflection direction and the observation direction, and calculate the overall main direction angle of the target electric meter image; Dividing the standardized electric meter image data into a plurality of image block sequences with overlapping areas in a two-dimensional space; The center position of each image block is converted into polar coordinates based on the image center point as the reference origin; Based on the polar coordinate expression, the distance component and the angle component of each image block are further mapped into a complex position vector to obtain a rotation-invariant position vector; Each image block P k The corresponding rotation invariant position vector RIPE k Splicing to form a sequence of image blocks with rotation-invariant position encoding.
4. The electric meter reading error correction method based on deep learning according to claim 1 is characterized in that: The improved Transformer encoder includes: The image block sequence with rotation-invariant position encoding is input into the improved Transformer encoder. The visual feature vector of each image block with rotation-invariant position encoding is concatenated with its rotation-invariant position vector. The local adaptive query vector is calculated through linear mapping and layer normalization. The visual similarity is calculated by comparing the visual feature vector with the visual feature vectors of all adjacent image blocks to generate the neighborhood mask weight. After concatenating the visual feature vector of each image block with the rotation-invariant position vector, a key vector and a value vector of the structural constraint are generated through a linear mapping function. Using locally adaptive query vectors Key vector and value vector with structural constraints Computational Improvement of Transformer Encoder Structure to Enhance Self-Attention Weights Where d is the dimension of the key vector, To improve the structure of the Transformer encoder, the self-attention weight is enhanced to indicate the degree to which the k-th image block pays attention to the structural information of the j-th image block during the encoding process; The structure-enhanced self-attention weights are used to perform weighted summation on the value vectors of the structural constraints of all image blocks to obtain the first encoded feature representation of the structure enhancement.
5. The electric meter reading error correction method based on deep learning according to claim 1 is characterized in that: The output three-dimensional cross-attention feature sequence includes: Reconstructing the structure-enhanced first encoded feature representation set into a three-dimensional feature tensor; In the spatial dimension, channel dimension, and depth or time dimension of the three-dimensional feature tensor, the query vector, key vector, and value vector of the spatial dimension, channel dimension, and depth or time dimension are calculated for each feature block respectively; Calculate the spatial attention weight A based on the query vector and key vector in the spatial dimension, channel dimension, and depth or time dimension s , channel attention weight A c and the depth or temporal attention weight A d ; Perform weighted aggregation on the attention weights to form a three-dimensional cross-attention output feature tensor F 3D : F 3D =a s ·(A s ·V s )+a c ·(A c ·V c )+a d ·(A d ·V d ); Among them, α s , α c , α d is the learnable weight factor in three dimensions: space, channel and depth, V s 、V c 、V d are value vectors of spatial dimension, channel dimension, and depth or time dimension respectively; The 3D cross-attention output feature tensor is flattened into a 3D cross-attention feature sequence according to the spatial arrangement order of the image blocks.
6. The electric meter reading error correction method based on deep learning according to claim 1 is characterized in that: The confidence assessment of the preliminary electric meter digital reading result is performed, comprising: For each feature vector in the three-dimensional cross-attention feature sequence, a set of decoder query vectors is calculated using linear transformation and layer normalization. The historical meter reading prior information sequence is transformed into a prior embedding sequence through an embedding matrix and sequentially concatenated with the decoder query vector set to form a joint decoding input sequence. The joint decoding input sequence is used to simultaneously introduce current features and historical prior information during the decoding process. In each layer of the Transformer decoder, the joint decoding input sequence is transformed into the value vector required for self-attention through the projection matrix, and the value vector is weighted and summed to obtain the output features of the decoding layer; Based on the output features of the last layer of decoder, a position-sensitive digital classification head is connected to generate a digital category logarithmic vector z for each sequence position. t =[z t,0 ,z t,1 ,…,z t,9 ], calculate the digital probability distribution by adjusting the Softmax function by temperature: Among them, T>0 is an adjustable temperature coefficient used to control the smoothness of the probability distribution; For each sequence position, the digital category with the maximum probability in the digital probability distribution is selected as the preliminary meter digital reading result, and the maximum probability is used as the confidence value of the current position; All confidence values in the confidence vector are averaged to obtain an overall confidence evaluation result. When the overall confidence evaluation result is greater than or equal to a preset confidence threshold, a high confidence evaluation result is output; otherwise, a low confidence evaluation result is output.
7. The method for correcting electric meter readings based on deep learning according to claim 1, characterized in that: The conditions for triggering the dual-modal enhanced reasoning mechanism include: When the overall confidence assessment result is less than the preset confidence threshold or the difference between the preliminary meter digital reading result and the historical meter reading prior information sequence is greater than the preset deviation range threshold, the dual-modal enhanced reasoning mechanism is automatically triggered.
8. The electric meter reading error correction method based on deep learning according to claim 7 is characterized in that: The dual-modal enhanced reasoning mechanism includes: For the standardized meter image data, the original perspective input and the corresponding mirror perspective input are constructed respectively; Calculate the original perspective enhanced reading results and the mirror perspective enhanced reading results of the original perspective input and the mirror perspective input, and calculate the corresponding confidence evaluation results respectively; Compare the confidence assessment results of the corresponding positions of the original perspective enhanced reading result and the mirror perspective enhanced reading result. If the confidence assessment result of the original perspective enhanced reading result is not less than the confidence assessment result of the mirror perspective enhanced reading result, the final enhanced meter digital reading result uses the original perspective enhanced reading result; otherwise, the mirror perspective enhanced reading result is used. The final enhanced meter digital reading result is the set of numbers at each position obtained by confidence comparison.
9. The electric meter reading error correction method based on deep learning according to claim 1, characterized in that: The calculation of the final meter digital reading result includes: Calculating the confidence value of each digit of the preliminary meter digital reading result and the enhanced meter digital reading result respectively; For each bit of the digital sequence, a number with a confidence value greater than a threshold is selected from the preliminary meter digital reading result and the enhanced meter digital reading result as the final meter digital reading result of that bit. The final meter digital reading result is composed of the predicted numbers with a confidence value greater than the threshold at each position.
Citation Information
Cited By
A time series prediction method and electronic device for cloud server cluster load scheduling
CN121561347B