Meter identification method based on deep learning and dynamic distortion correction

Through deep learning and dynamic distortion correction methods, combined with multi-scale fusion object detection and key point detection framework, the distortion problem of meter recognition in robot inspection is solved, and the accuracy of instrument readings and network adaptability is improved.

CN120544211AInactive Publication Date: 2025-08-26浙江浙能数字科技有限公司

Patent Information

Application Number
CN202511036720.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing meter recognition technology reduces the recognition accuracy due to changes in target proportion and perspective changes in robot inspection, especially in the case of distortion caused by robot positioning errors, and the template matching method fails.

Method used

The multi-scale fusion object detection framework based on deep learning is combined with the key point detection framework, and image distortion is corrected through feature point matching and perspective transformation, the RFV-SPPF module is used to enhance multi-scale perception capabilities, the Ghost module is used to reduce the calculation amount, and the abnormal points are eliminated through SIFT and RANSAC algorithms, and the homography matrix is ​​calculated for image correction.

Benefits of technology

It improves the calculation accuracy of instrument readings, enhances the robustness and adaptability of the network, and is suitable for complex power plant robot inspection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544211A_ABST
    Figure CN120544211A_ABST
Patent Text Reader

Abstract

The invention relates to a meter identification method based on deep learning and dynamic distortion correction, and the method comprises the steps: obtaining a source image which comprises a dial region; positioning and extracting a dial region in the source image by using a multi-scale fused target detection frame; feature point matching is carried out on the zoomed image and a template image, and abnormal points are removed; calculating a homography matrix by using the corresponding relation of the feature points; identifying key points in the corrected image through a key point detection frame; and calculating the meter reading according to the geometrical relationship of the key points. The method has the beneficial effects that the homography matrix is calculated through feature point extraction and abnormal point elimination, and the problem of dial target distortion caused by a visual angle in the robot inspection process is solved by using a perspective transformation mode. Moreover, the instrument is approximately restored into a front-view image by correcting distortion, so that the calculation precision of the reading of the instrument can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition technology, and in particular relates to a meter recognition method based on deep learning and dynamic distortion correction. Background Art

[0002] Existing meter recognition technologies mostly use image template matching, that is, using predefined templates to locate and read the dial in the actual scene. Figure 1 As shown, the dial in the upper left corner is the template, and the one in the lower right corner is the source image. By sliding the template from left to right and from top to bottom, the similarity between the template and the image block is calculated at each position. The maximum similarity is achieved when the dial in the template image completely overlaps with the dial in the source image. However, this method has many limitations. For example, template matching may fail when the target changes in scale or perspective. During robot inspections, robot positioning errors caused by odometry accuracy may cause changes in the scale and perspective of the dial target in the real-time image. Template matching algorithms rely on templates of fixed size, so a more accurate and versatile dial recognition method is needed to cope with complex application scenarios. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a meter recognition method based on deep learning and dynamic distortion correction.

[0004] First, a meter recognition method based on deep learning and dynamic distortion correction is provided, including:

[0005] Step 1: Acquire a source image, where the source image includes a dial area;

[0006] Step 2: Using a multi-scale fusion object detection framework, locate and extract the dial area in the source image; and scale the dial area to the same size as the pre-stored template image;

[0007] Step 3: Match the feature points of the scaled image with the template image and remove abnormal points;

[0008] Step 4: Calculate the homography matrix using the correspondence between the feature points; and obtain the corrected image based on the homography matrix;

[0009] Step 5: Identify key points in the corrected image using a key point detection framework, where the key points include the starting point, the ending point, the center point of the dial, and the tip of the hands.

[0010] Step 6: Calculate the instrument readings based on the geometric relationship of the key points.

[0011] Preferably, in step 2, the target detection framework is an improved YOLOv8 network, and its SPPF module is replaced by an RFV-SPPF module; the RFV-SPPF module includes a variety of convolution kernel combinations.

[0012] Preferably, in step 5, the key point detection framework includes a Ghost module, and the Ghost module performs the following operations:

[0013] Perform channel compression convolution on the input feature map to generate the first feature map;

[0014] Performing a secondary convolution on the first feature map to generate a second feature map;

[0015] Concatenate the first feature map and the second feature map along the channel dimension, and output a feature map of the same dimension as the input.

[0016] Preferably, in step 6, the angle between the pointer and the starting point of the range, as well as the proportion of the pointer in the total range of the range are determined by geometric calculation, and then the instrument reading is calculated.

[0017] Preferably, in step 3, feature point matching is performed using a SIFT algorithm, a SURF algorithm, or a GPU-based ORB algorithm; and outlier removal is performed using a RANSAC algorithm or KNN.

[0018] In a second aspect, a meter recognition system based on deep learning and dynamic distortion correction is provided, which is used to execute any of the methods described in the first aspect, including:

[0019] An acquisition module, configured to acquire a source image, wherein the source image includes a dial area;

[0020] A positioning module is used to locate and extract the dial area in the source image using a multi-scale fusion object detection framework; and scale the dial area to the same size as the pre-stored template image;

[0021] The matching module is used to match the feature points of the scaled image with the template image and remove abnormal points;

[0022] A first calculation module is used to calculate a homography matrix using the correspondence between the feature points; and obtain a corrected image according to the homography matrix;

[0023] an identification module for identifying key points in the corrected image using a key point detection framework, wherein the key points include a starting point, an end point, a center point of the dial, and a tip position of the hands;

[0024] The second calculation module is used to calculate the instrument reading according to the geometric relationship of the key points.

[0025] According to a third aspect, a computer storage medium is provided, wherein a computer program is stored in the computer storage medium; when the computer program is executed on a computer, the computer executes any one of the methods described in the first aspect.

[0026] In a fourth aspect, an electronic device is provided, including:

[0027] Memory, used to store computer programs;

[0028] A processor is used to execute the computer program to implement any method as described in the first aspect.

[0029] The beneficial effects of the present invention are:

[0030] 1. This invention calculates a homography matrix by extracting feature points and eliminating outliers, using perspective transformation to address the issue of dial distortion caused by viewing angles during robotic inspections. Furthermore, by correcting the distortion and restoring the instrument to a near-normal image, the accuracy of the instrument readings can be improved.

[0031] 2. This paper trains and deploys an end-to-end target detection framework that supports multi-scale feature fusion for instrument target localization, enhancing network robustness and enabling target instrument localization in complex power plant robotic inspection scenarios. Furthermore, this paper increases the receptive field by using different convolution kernel sizes in RFV-SPPF, enhancing the network's multi-scale perception capabilities, and employs shortcut residual connections to optimize gradient computation.

[0032] 3. The present invention trains and deploys a lightweight end-to-end key point detection framework for both ends of the instrument range and instrument needle positioning. Using the Ghost module to replace the traditional convolution can reduce the computational complexity of the network and optimize the deployment of mobile robots or edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the template matching algorithm provided by the present invention;

[0034] Figure 2 This is a meter identification flow chart provided by the present invention;

[0035] Figure 3 This is a schematic diagram of the structure of the target detection framework used in the present invention;

[0036] Figure 4 A schematic diagram comparing the SPPF module and the RFV-SPPF module provided by the present invention;

[0037] Figure 5 A schematic diagram of the Ghost module structure provided by the present invention;

[0038] Figure 6 This is a schematic diagram of the impact of image distortion correction on instrument reading recognition provided by the present invention. DETAILED DESCRIPTION

[0039] The present invention will be further described below with reference to the following examples. The following examples are provided only to facilitate understanding of the present invention. It should be noted that, without departing from the principles of the present invention, it is possible for a person skilled in the art to make various modifications to the present invention, and such improvements and modifications fall within the scope of the claims of the present invention.

[0040] Example 1:

[0041] To solve the problems of the prior art, Example 1 of the present application provides a meter recognition method based on deep learning and dynamic distortion correction, such as Figure 2 Shown, including:

[0042] Step 1: Acquire a source image, where the source image includes a dial area.

[0043] Specifically, the source image is the actual scene picture taken by the inspection robot, such as Figure 2 As shown in the image, the dial is distorted; the originally round dial appears elliptical due to the shooting angle. This distortion poses a challenge to conventional template matching-based methods and deep learning-based methods, making it difficult for these methods to accurately recognize the dial or resulting in low recognition accuracy.

[0044] In this paper, we propose a solution: combining a deep learning-based target detection framework with a key point detection framework, and incorporating digital image processing techniques (such as SIFT and RANSAC), to perform perspective transformation on the instrument target, thereby obtaining an approximately orthographic image (for details, please refer to subsequent steps 2-4).

[0045] Step 2: Use a multi-scale fusion target detection framework to locate and extract the dial area in the source image; and scale the dial area to the same size as the pre-stored template image.

[0046] Specifically, a multi-scale fusion object detection framework is used to locate and identify target instruments in the image. After the target is identified, it is extracted from the source image and scaled proportionally to the template image to match its aspect ratio.

[0047] Among them, such as Figure 3As shown in the figure, the target detection framework is an improved YOLOv8 network. The traditional SPPF module is modified into the RFV-SPPF (Receptive Field Variation SPPF, receptive field variation spatial pyramid pooling - fast) module. During the convolution process, different receptive field sizes are added to enhance the multi-scale perception capability of the network.

[0048] Figure 4 To compare SPPF and RFV-SPPF, the SPPF module uses three consecutive MaxPool2d operations for multi-scale feature extraction, aggregates features at different pooling scales through Concat, and uses Conv + BN + SiLU activation to enhance expressiveness. The RFV-SPPF module, on the other hand, retains the multi-scale pooling path of SPPF and introduces the RFV block (receptive field transformation module) to replace the original convolution operation. The RFV block contains various convolution kernel combinations (such as 1×1, 3×3, and 5×5) to simulate receptive fields of different sizes. It uses shortcut residual connections and SiLU activation to improve expressiveness and gradient propagation capabilities compared to SPPF.

[0049] In addition, the scaling process is:

[0050] The size of the target detection frame output by the target detection framework is ( ), and the size of the template image is ( In order to perform template matching or feature comparison, the image area captured by the target detection frame is scaled to the same size as the template image ( ), thereby ensuring input consistency and matching accuracy.

[0051] Step 3: Match the feature points of the scaled image with the template image and remove abnormal points.

[0052] In step 3, feature point matching is performed using the SIFT algorithm, SURF algorithm, or GPU-based ORB algorithm; outliers are removed using the RANSAC algorithm or KNN.

[0053] Step 4: Calculate the homography matrix using the correspondence between the feature points; and obtain the corrected image based on the homography matrix.

[0054] Step 5: Identify key points in the corrected image using a key point detection framework, where the key points include a starting point, an end point, the center point of the dial, and the tip position of the hands.

[0055] Step 6: Calculate the instrument readings based on the geometric relationship of the key points.

[0056] Example 2:

[0057] Based on Example 1, Example 2 of the present application provides a more specific meter recognition method based on deep learning and dynamic distortion correction, including:

[0058] Step 1: Acquire a source image, where the source image includes a dial area.

[0059] Step 2: Use a multi-scale fusion target detection framework to locate and extract the dial area in the source image; and scale the dial area to the same size as the pre-stored template image.

[0060] Step 3: Match the feature points of the scaled image with the template image and remove abnormal points.

[0061] For example, the present embodiment uses the SIFT (Scale-Invariant Feature Transform) algorithm to match feature points between the scaled image and the template image, ensuring a one-to-one correspondence between the same feature points in the two images. Subsequently, the RANSAC (Random Sample Consensus) algorithm is used to remove abnormal feature points generated during the matching process, ensuring the accuracy of feature point matching.

[0062] Following are the steps of SIFT algorithm:

[0063] (1) Gaussian scale space generation

[0064] Use Gaussian blur to construct the scale space,

[0065]

[0066] The Gaussian kernel for , is the input image, Used to control the degree of blur.

[0067] (2) Gaussian difference image

[0068] It is obtained by subtracting adjacent scale images.

[0069]

[0070] in is the scale coefficient.

[0071] (3) Feature point positioning

[0072] In the Gaussian difference image, the extreme points are fitted by Taylor expansion.

[0073]

[0074] The offset .

[0075] (4) Feature point direction allocation

[0076] The direction histogram is constructed by calculating the gradient amplitude and direction, and the main direction is the direction of the histogram peak.

[0077]

[0078]

[0079] (5) Feature point description

[0080] Taking the feature point as the center, divide Field, statistics for each The gradient direction histogram of the sub-region generates a 128-dimensional feature vector,

[0081]

[0082] (6) Feature point matching

[0083] The similarity is calculated by Euclidean distance.

[0084]

[0085] Use Lowe ratio test to filter matching points. If , then the match is successful, where and are the nearest neighbor and next nearest neighbor distances, respectively.

[0086] In addition, the present embodiment uses the RANSAC algorithm to remove outliers from the feature points obtained by SIFT matching. The RANSAC algorithm is an iterative algorithm for estimating mathematical model parameters from a dataset containing outliers (noise or erroneous data points). It robustly finds the optimal model in uncertain data through random sampling and model verification. The RANSAC algorithm steps are as follows:

[0087] (1) Initialization

[0088] Set the number of iterations N, the inlier threshold t, and the minimum number of sampling points s.

[0089] (2) Iteration

[0090] Randomly sample s data points, fit the model according to the sampled points, calculate the error of all points to the model, count the number of inliers with an error threshold of t, and update the optimal model if the number of inliers exceeds the current model.

[0091] (3) Termination conditions are met

[0092] The maximum number of iterations N is reached or a model that meets the requirements is found.

[0093] Step 4: Calculate the homography matrix using the correspondence between the feature points; and obtain the corrected image based on the homography matrix.

[0094] The role of the homography matrix is ​​to perform a geometric transformation on the image, correcting the image's perspective distortion, thereby obtaining an image that is approximately orthographic. This application uses SIFT and RANSAC to obtain matching feature points and then calculate the homography matrix H, thereby obtaining an image that is approximately orthographic after the perspective transformation.

[0095] Step 5: Identify key points in the corrected image using a key point detection framework, where the key points include a starting point, an end point, the center point of the dial, and the tip position of the hands.

[0096] Key point detection framework uses Figure 3 The key point detection network in the paper is based on YOLOv8, but the network parameters are lightweight. The present invention uses the Ghost module, which can replace the convolution module in the YOLOv8 network without changing the dimension of the input feature map, thereby reducing the network calculation amount.

[0097] For example, Figure 5 As shown in the figure, the Ghost module uses the concept of identity mapping, preserving the original feature map and adding it to the convolved feature map along the channel dimension. Specifically, assuming the input feature map has dimensions (64, 64, 6), a convolution operation is first performed to compress the channel dimension, resulting in a first feature map with dimensions (64, 64, 3). Subsequently, a convolution operation is applied to the first feature map again for feature extraction, outputting a second feature map with dimensions (64, 64, 3). Both the first and second feature maps serve as intermediate feature maps. Finally, this feature-extracted second feature map is stacked with the first feature map along the channel dimension, resulting in an output feature map with dimensions (64, 64, 6), which matches the input feature map.

[0098] It should be noted that the target detection framework and key point detection framework can also use other mainstream CNN networks such as fast-RCNN or YOLO models other than YOLOv8.

[0099] Step 6: Calculate the instrument readings based on the geometric relationship of the key points.

[0100] In step 6, after obtaining the coordinates of the key points, geometric calculations are used to determine the angle between the needle and the starting point of the range, as well as the proportion of the needle in the total range. Based on this information, the meter reading can be calculated.

[0101] like Figure 6 As shown in the figure, after extensive testing in actual robot inspection scenarios in power plants, the actual usage scenarios included more than 10 different instruments. The error when using image distortion correction was much smaller than the error when not applying image distortion correction.

[0102] It should be noted that the parts in this embodiment that are the same or similar to those in Example 1 can be referenced to each other and will not be described in detail in this application.

[0103] Example 3:

[0104] Based on Example 2, Example 3 of the present application provides a meter recognition system based on deep learning and dynamic distortion correction, including:

[0105] An acquisition module, configured to acquire a source image, wherein the source image includes a dial area;

[0106] A positioning module is used to locate and extract the dial area in the source image using a multi-scale fusion object detection framework; and scale the dial area to the same size as the pre-stored template image;

[0107] The matching module is used to match the feature points of the scaled image with the template image and remove abnormal points;

[0108] A first calculation module is used to calculate a homography matrix using the correspondence between the feature points; and obtain a corrected image according to the homography matrix;

[0109] an identification module for identifying key points in the corrected image using a key point detection framework, wherein the key points include a starting point, an end point, a center point of the dial, and a tip position of the hands;

[0110] The second calculation module is used to calculate the instrument reading according to the geometric relationship of the key points.

[0111] It should be noted that the system provided in this embodiment is a system corresponding to the method provided in Example 2. Therefore, the parts in this embodiment that are the same or similar to those in Example 2 can be referenced to each other and will not be repeated in this application.

[0112] In summary, this application trains and deploys an end-to-end target detection framework that supports multi-scale feature fusion and a lightweight end-to-end key point detection framework for instrument target positioning, instrument range ends, and instrument needle positioning, and calculates the homography matrix through feature point extraction and outlier removal. The perspective transformation method is used to solve the problem of dial target distortion caused by viewing angle during robot inspection. It can improve the accuracy of existing dial counting and recognition, and has good versatility in robot inspection scenarios (applicable to various different industrial instruments).

Claims

1. A meter recognition method based on deep learning and dynamic distortion correction, characterized in that: include: Step 1: Acquire a source image, where the source image includes a dial area; Step 2: Use the multi-scale fusion object detection framework to locate and extract the dial area in the source image; and scaling the dial area to the same size as the pre-stored template image; Step 3: Match the feature points of the scaled image with the template image and remove abnormal points; Step 4: Calculate the homography matrix using the correspondence between the feature points; and obtain the corrected image based on the homography matrix; Step 5: Identify key points in the corrected image using a key point detection framework, where the key points include the starting point, the ending point, the center point of the dial, and the tip of the hands. Step 6: Calculate the instrument readings based on the geometric relationship of the key points.

2. The meter recognition method based on deep learning and dynamic distortion correction according to claim 1, characterized in that: In step 2, the target detection framework is an improved YOLOv8 network, and its SPPF module is replaced by an RFV-SPPF module; the RFV-SPPF module contains multiple convolution kernel combinations.

3. The meter recognition method based on deep learning and dynamic distortion correction according to claim 2, characterized in that: In step 5, the key point detection framework includes a Ghost module, which performs the following operations: Perform channel compression convolution on the input feature map to generate the first feature map; Performing a secondary convolution on the first feature map to generate a second feature map; Concatenate the first feature map and the second feature map along the channel dimension, and output a feature map of the same dimension as the input.

4. The meter recognition method based on deep learning and dynamic distortion correction according to claim 3, characterized in that: In step 6, geometric calculations are used to determine the angle between the meter needle and the starting point of the range, as well as the proportion of the meter needle in the total range, and then the meter reading is calculated.

5. The meter recognition method based on deep learning and dynamic distortion correction according to claim 4, characterized in that: In step 3, feature point matching is performed using the SIFT algorithm, SURF algorithm, or GPU-based ORB algorithm; outliers are removed using the RANSAC algorithm or KNN.

6. A meter recognition system based on deep learning and dynamic distortion correction, characterized by: Used to perform the method according to any one of claims 1 to 5, comprising: An acquisition module, configured to acquire a source image, wherein the source image includes a dial area; A positioning module is used to locate and extract the dial area in the source image using a multi-scale fusion object detection framework; and scale the dial area to the same size as the pre-stored template image; The matching module is used to match the feature points of the scaled image with the template image and remove abnormal points; A first calculation module is used to calculate a homography matrix using the correspondence between the feature points; and obtain a corrected image according to the homography matrix; an identification module for identifying key points in the corrected image using a key point detection framework, wherein the key points include a starting point, an end point, a center point of the dial, and a tip position of the hands; The second calculation module is used to calculate the instrument reading according to the geometric relationship of the key points.

7. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is run on a computer, the computer executes the method according to any one of claims 1 to 5.

8. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Instrument image recognition method and device

    CN109145699A

  • Pointer meter positioning and reading algorithm based on dial plate feature

    CN110111387A

  • Pointer type instrument image tilt correction method

    CN112801094A

  • Improved YOLOv8-based real-time bearing defect detection method

    CN117788428A

  • Light-weight fabric defect detection algorithm based on deformable convolution and high-level screening feature pyramid

    CN119963470A

Cited By

  • Meter identification method based on calibration and relocation

    CN121708574A