Instrument reading identification method based on self-generation and annotation specification benchmark test set
By using a method of self-generating and annotating standardized benchmark test sets, and generating instrument datasets using Python and Blender, combined with the DINOv2 model and RANSAC algorithm, the problems of data scarcity and high annotation costs in automated instrument reading algorithms are solved, achieving high-precision automated reading recognition.
Patent Information
- Application Number
- CN202511105825.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-07
AI Technical Summary
In existing technologies, automated reading algorithms for analog instruments suffer from data scarcity and high annotation costs, resulting in insufficient reading recognition accuracy, requiring manual verification, and lacking effective high-quality annotated datasets.
We adopted a method of self-generated and annotated standard benchmark test sets, used Python scripts and Blender to generate instrument datasets, combined the DINOv2 heat map prediction model and key point extraction algorithm for scale marking and pointer prediction, and used a regression network for perspective correction. We also used the RANSAC algorithm to fit the relationship between polar angle and gauge reading.
It effectively solves the problem of scarce standardized instrument datasets, improves the accuracy of instrument reading recognition, reduces the need for manual verification, and achieves high-precision automated reading recognition.
Smart Images

Figure CN120997813A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of instrument transmission technology, and more specifically to an instrument reading recognition method based on a self-generated and annotated specification benchmark test set. Background Technology
[0002] In industrial automation systems, many devices transmit critical data via analog instruments. These instruments are used not only to detect system anomalies but also to maintain operational stability through the control system. Traditional manual reading methods are inefficient and particularly impractical in large-scale industrial applications. However, with advancements in industrial inspection equipment, including straightness testers, drones, and robotic platforms integrating computer vision algorithms, the instrument reading process has gradually become automated. Despite these systems achieving a high level of automation, limitations remain in reading accuracy, necessitating secondary manual verification. This challenge is particularly pronounced in the reading of analog instruments, as it requires comprehensive consideration of multiple visual elements (such as dials, scale lines, pointers, and instrument numbers) to deduce the measured value, placing stringent demands on the reading algorithm.
[0003] The scarcity of standardized instrument datasets is a major obstacle to the development of algorithms in this field, and the development of analog instrument reading algorithms is significantly limited by data quality issues. In addition, there is a lack of high-quality annotations required for effective algorithm training. The three main factors causing data scarcity are: (1) Data acquisition challenges: Under stable operating conditions, the changes in the measured physical quantities are negligible, resulting in minimal changes in the captured instrument images. In contrast, manually taken photos are not only labor-intensive but also time-consuming, which hinders the effective creation of large-scale datasets. (2) High annotation costs: State-of-the-art instrument reading algorithms typically employ deep learning architectures, decomposing the task into multiple subtasks, such as instrument detection, pointer recognition, and scale mark recognition, each of which requires specialized annotation. Some implementations also require semantic segmentation or keypoint detection of scale marks, which necessitates meticulous annotation of many fine-grained features in a single image. (3) Data privacy restrictions: In industrial environments, instrument systems are embedded in complex operating frameworks and are subject to strict privacy regulations, which greatly limits access to third-party data.
[0004] Meanwhile, the algorithm for measuring and reading instrument data needs further development.
[0005] Therefore, how to provide an instrument reading recognition method based on a self-generated and annotated standard benchmark test set that can effectively solve the obstacle to the development of algorithms in this field caused by the scarcity of standardized instrument datasets, and further develop high-precision instrument data measurement and reading algorithms, and effectively improve the accuracy of instrument reading recognition, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] In view of this, the present invention provides a method for instrument reading identification based on a self-generated and annotated standard benchmark test set.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A method for instrument reading identification based on a self-generated and annotated specification benchmark test set includes:
[0009] Step 1: Use a Python script to randomly sample and set the measurement and appearance parameters, and use Blender to generate the instrument dataset;
[0010] Step 2: Construct a fully automated specification annotation framework based on Blender 3D model parameters to annotate the instrument dataset and obtain the specification benchmark test set;
[0011] Step 3: Scale marking and pointer prediction are performed using the DINOv2-based heatmap prediction model and key point extraction algorithm. In the key point extraction, a multi-stage extraction method is used for intermediate scale marking that may contain an uncertain number of key points.
[0012] Step 4: After text detection and recognition, the homography matrix required for perspective correction is directly regressed from the original image end-to-end using a regression network. The relationship between the polar angle and the gauge reading is fitted using the RANSAC algorithm, and the value corresponding to the pointer polar angle is predicted through the fitted model.
[0013] Optionally, in step 1, the measurement parameters include: the rotation angle of the pointer and the measurement range of the instrument.
[0014] Optionally, in step 1, appearance parameters include: the color and shape of the pointer, the length of the scale marks and the numerical spacing between adjacent marks, the material and corrosion condition of the instrument housing, the angle of the camera and its distance from the instrument, the style and size of the digital text and its radial position relative to the center, the texture and noise on the dial, and the diversity of the environment.
[0015] Optionally, in step 2, the fully automated specification annotation framework based on Blender 3D model parameters is as follows:
[0016] Gauge inspection and annotation of the bounding box: In Blender, traverse all vertices of the gauge shell and calculate the projected coordinates of the vertices in the camera image; convert the projected coordinates into a two-dimensional coordinate system with the top left corner as the origin, and summarize the maximum and minimum xy values of the projected coordinates to determine the bounding box of the gauge shell;
[0017] Perspective correction: Establish the correspondence between two point sets: the positions of four cocircular equidistant points in the world coordinate system and the corresponding points of the camera projection of the four cocircular equidistant points; normalize both coordinate sets relative to the detected scale bounding boxes, and use OpenCV's isomorphic estimation method to calculate the perspective transformation matrix between the normalized coordinate systems, which serves as the training target for the distortion correction model.
[0018] Scale mark detection: divided into key point detection and semantic segmentation; for key point detection, points adjacent to the center of the gauge are marked at each major scale mark; for semantic segmentation, polygon annotation is performed based on the geometry of the scale mark.
[0019] Point detection is divided into key point detection and semantic segmentation. For key point detection, the intersection of the pointer rotation axis and the rotation axis surface is defined as the starting position, and the ending position is determined by traversing all points on the pointer and selecting the point farthest from the starting position. For pointer segmentation and annotation, a mask-based annotation method is used.
[0020] Optical Character Recognition (OCR) technology includes text localization and text recognition. It generates bounding boxes for all text objects and associates the text content corresponding to the bounding boxes for recognition.
[0021] Optionally, step 3, before using the DINOv2-based heatmap prediction model and key point extraction algorithm for scale marking and pointer prediction, also includes: using YOLOv11s to locate the instrument and separate the instrument from the complex background.
[0022] Optionally, in step 3, scale marking and pointer prediction are performed using a DINOv2-based heatmap prediction model and keypoint extraction algorithm, specifically as follows:
[0023] Five keypoint categories were selected, including: start scale marker, end scale marker, intermediate scale marker, pointer start position, and pointer end position;
[0024] Heatmap prediction models are used to predict heatmap values for each keypoint category. The heatmap prediction model is as follows:
[0025] The frozen pre-trained DINOv2 is used as an encoder to extract features from the input image; a multi-stage upsampling network is used to generate a heatmap of the same size as the original image, and BCE loss is used as the prediction loss; each upsampling stage consists of a convolutional layer and a transposed convolutional layer, the latter performing 2× upsampling, and the final heatmap is generated by a convolutional layer.
[0026] Keypoints are extracted from the heatmaps of each keypoint category. For intermediate scale labels, each heatmap corresponds to one keypoint, and the mean shift algorithm is used for extraction.
[0027] Optionally, in step 3, a multi-stage extraction method is used for intermediate-scale markers that may contain an uncertain number of keypoints during keypoint extraction. Specifically:
[0028] Significant region filtering: Set a response threshold c, and retain only candidate regions whose heatmap responses satisfy H(x,y)≥c;
[0029] Local maximum extraction: A k×k sliding window is used to perform maximum filtering to detect local maxima that satisfy the following conditions:
[0030]
[0031] Non-maximum suppression: A greedy algorithm based on spatial distance constraints is adopted to iteratively select key points in descending order of confidence, and low-confidence candidate points within the radius r of the selected key points are eliminated;
[0032] Subpixel thinning: Integer pixel coordinates are adjusted using quadratic interpolation, and peak offset is estimated using Taylor expansion.
[0033]
[0034] The keypoint coordinates (x, y) are optimized to (x+Δx, y+Δy), where both |Δx| and |Δy| do not exceed 0.5.
[0035] Optionally, in step 4, YOLOv11s is used for text detection and ABINet is used for text recognition.
[0036] Optionally, in step 4, a regression network is used to directly regress the homography matrix required for perspective correction from end to end in the original image to perform perspective correction, specifically as follows:
[0037] A deep learning model is trained using homography annotations in synthetic data to predict homography transformations;
[0038] Homography transformation is represented by eight degrees of freedom, and the model predicts a vector h∈R. 1×8 The pre-trained DINO backbone network is used as the encoder to extract features from the input image and predict the final homography transformation through a linear layer, and is trained using the L1 loss function.
[0039] During inference, a 1 is added to the end of the prediction vector h, and it is reshaped into an isomorphic matrix suitable for images with normalized width and height. For an image of shape (h, w), the isomorphic transformation matrix H homo Calculate as follows:
[0040]
[0041] After obtaining the isomorphic transformation, it is applied to the key points and text coordinates detected in the previous stages and mapped onto the corrected perspective plane.
[0042] Optionally, in step 4, the process of obtaining the relationship between the polar angle and the gauge reading is as follows:
[0043] The OCR output is filtered based on confidence scores by combining detection and recognition.
[0044] A polar coordinate system was established with the starting position of the pointer as the pole and the downward vertical direction as the polar axis. The polar angles of the scale mark, the pointer endpoint, and the center of the OCR bounding box were calculated.
[0045] Establish the relationship between the polar angle and the gauge reading, when the polar angle marked on the ruler... Polar angle with the center of the OCR bounding box The two will be matched if the following conditions are met:
[0046]
[0047] min(d ij ,(360-d ij ))≤d;
[0048] Where d is the threshold; when two or more matches are detected, the correspondence between the polar angles of these scale markers and their numerical values is recorded.
[0049] As can be seen from the above technical solution, compared with the prior art, this invention discloses a method for instrument reading recognition based on a self-generated and annotated standard benchmark test set. It randomly samples measurement and appearance parameters using Python scripts, and generates an instrument dataset using Blender. A standard benchmark test set is obtained using a fully automatic standard annotation framework based on Blender 3D model parameters. Then, a DINOv2-based heatmap prediction model and a keypoint extraction algorithm (including a multi-stage extraction method) are used for scale marking and pointer prediction. Finally, after text detection and recognition, perspective correction is achieved by regressing the homography matrix using a regression network. The RANSAC algorithm is then used to fit the relationship between the polar angle and the gauge reading to predict the value. This effectively solves the problems of the scarcity of standardized instrument datasets hindering algorithm development and the insufficient accuracy of existing reading algorithms requiring secondary manual verification. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0051] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example 1:
[0054] Embodiment 1 of this invention discloses a method for instrument reading identification based on a self-generated and annotated specification benchmark test set, such as... Figure 1 As shown, it includes:
[0055] Step 1: Use a Python script to randomly sample and set the measurement and appearance parameters, and use Blender to generate the instrument dataset.
[0056] To address the limitations of instrument data in existing technologies, this invention proposes a synthetic analogy generation framework that utilizes Blender simulation software for parameterized control.
[0057] A well-designed instrument synthesis framework should possess two core characteristics: parameter control and realistic reproduction. Parameter control enables the synthesized instrument to be comprehensively adjusted through configurable parameters covering all aspects from visual attributes to operational range, thereby realistically simulating real-world instruments. Realistic reproduction requires not only realistic visual effects but also the systematic simulation of environmental interference factors, such as reflections from glass surfaces, obstructions caused by contaminants, and dynamic shadow patterns.
[0058] Based on this concept, this invention utilizes Blender's modeling capabilities to construct an instrument using six modular components: an instrument lens, pointer, instrument digits, scale markings, dial, and bezel. Through system decomposition, external parameters controlling material properties and spatial configuration are established, including measurement parameters and appearance parameters. Measurement parameters affect the instrument's numerical readings, including the pointer's rotation angle and the instrument's measurement range. Appearance parameters enhance the diversity of generated images by altering the appearance of instrument elements, environmental conditions, and camera perspective. These parameters include: the pointer's color and shape, the length of the scale markings and the numerical spacing between adjacent markings, the material and corrosion condition of the instrument casing, the camera angle and distance from the instrument, the style and size of the numerical text and its radial position relative to the center, texture and noise on the dial, and environmental diversity.
[0059] Hands: The two most crucial factors for hands are color and shape. This invention pre-designs a variety of hand styles, each differing in length, thickness, and outline. To create a striking contrast with the dial, the hands are defaulted to black or white, while additional color variations are introduced with a certain probability by randomly adjusting the RGB properties of the hand material.
[0060] Scale markings: Scale markings are parameterized by length and numerical spacing between adjacent markings, enabling the generation of various measurement ranges and spatial distributions. Layout patterns include radial alignment variations and arc connection configurations; these randomly selected layout patterns are designed to challenge recognition models, although their impact on human readability is negligible.
[0061] Instrument Housing: The material properties of the instrument housing vary in color, gloss, and texture, all of which can affect detection performance. To address these issues, this invention randomly assigns different materials to the instrument housing. Furthermore, rust textures are introduced to simulate the corrosion of metal housings in real-world environments.
[0062] Camera: The camera position is randomized to simulate changes in viewing angle while maintaining a 30° deviation from the surface normal. Furthermore, the distance between the camera and the instrument is also randomized to simulate different shooting distances.
[0063] Text Diversity: The shape of numbers varies across different fonts. The most significant difference lies in serif and sans-serif fonts; serif fonts feature decorative strokes at the beginning and end of characters, while sans-serif fonts have a simpler design. Furthermore, some fonts appear slender, while others are more rounded. These typographical differences affect subsequent OCR performance. To enhance diversity, this invention introduces multiple font styles when generating the gauges. Additionally, the size of the numerical text and its radial position relative to the center are parameterized, further increasing text diversity.
[0064] Dial: Many real-world instruments include text information on their dials to indicate the instrument type and unit of measurement. For example, an oil level dial might display "oil" and "%". To simulate various real-world dial types, this invention adjusts the dial texture accordingly. Furthermore, random noise of different colors, shapes, positions, and sizes is superimposed on the dial to simulate effects such as rust and dust accumulation.
[0065] Environmental Diversity: This invention collects a large amount of environmental textures, including indoor environments such as apartments, factories, and bathrooms, as well as outdoor environments such as grasslands, mountains, and lakes. When generating each scale model, an environment is randomly selected to provide a unique background for each synthetic image, thereby enhancing the robustness of subsequent scale detection models. Different environments also introduce diverse lighting conditions, including strong sunlight, soft indoor lighting, and dim ambient light, such as the faint light under a bridge. These settings result in diverse lighting conditions in terms of intensity, distance, and light source type (e.g., point light, area light). Utilizing Blender's optical rendering capabilities, this invention generates rich optical effects. Through cameras at different angles, these light sources produce varying degrees of refraction on the instrument's glass cover, simulating real-world reflections. Furthermore, the interaction of these lights projects shadows onto the dial and the instrument casing, creating more complex image features.
[0066] To achieve random sampling of measurement and appearance parameters, this invention utilizes Python scripts to automate the generation process and enhance output diversity. Through Blender's Python API, this invention implements instrument model generation, scene integration, camera transformation, and rendering export functions.
[0067] Step 2: Construct a fully automated specification annotation framework based on Blender 3D model parameters to annotate the instrument dataset and obtain the specification benchmark test set.
[0068] The main advantage of synthetic instruments lies in their automatic annotation function. Spatial coordinates of instrument components can be directly obtained from 3D models, and data extraction is automated through Python scripts. For typical instrument reading models, annotation is required at the following stages: instrument detection, viewpoint correction, scale mark detection, pointer detection, and OCR recognition. This invention lists the required annotations for each stage based on the relevant algorithm type.
[0069] A fully automated specification annotation framework based on Blender 3D model parameters, specifically:
[0070] Gauge detection and bounding box annotation: Using an object detection algorithm, it is necessary to annotate the bounding box of the gauge shell. In the Blender environment, each object is a geometric entity composed of vertices and edges. In Blender, all vertices of the gauge shell are traversed, and the projected coordinates of the vertices in the camera image are calculated. The projected coordinates are then converted to a two-dimensional coordinate system with the top-left corner as the origin. The maximum and minimum x and y values of the projected coordinates are summed to determine the bounding box of the gauge shell.
[0071] Perspective Correction: Virtual instruments offer inherent advantages for perspective correction by directly accessing real spatial relationships. This invention establishes a correspondence between two point sets: the positions of four equidistant points on a circular scale in the world coordinate system (representing the frontal view) and the corresponding points of the camera projections of these four equidistant points. To focus on the perspective distortion of a specific scale, both coordinate sets are normalized relative to the detected scale bounding box. Specifically, the four sampling points on the scale and their projected coordinates in the camera view are transformed into a coordinate system with the top-left corner of the detected scale bounding box as the origin. These coordinates are then normalized according to the width and height of the bounding box. Finally, the perspective transformation matrix between the normalized coordinate systems is calculated using OpenCV's isomorphism estimation method, serving as the training target for the distortion correction model.
[0072] Scale mark detection: It is divided into two main algorithm types: key point detection and semantic segmentation. For key point detection, the point adjacent to the center of the gauge is marked at each major scale mark. For semantic segmentation, polygon labeling is performed according to the geometry of the scale mark, that is, the outermost vertex is selected sequentially along the outline of each scale mark to construct a closed polygon.
[0073] Point detection: Similar to scale marker detection, it is divided into keypoint detection and semantic segmentation. For keypoint detection, the intersection of the pointer rotation axis and the rotation axis surface is defined as the starting position, and the ending position is determined by traversing all points on the pointer and selecting the point farthest from the starting position. For pointer segmentation annotation, a mask-based annotation method is used. Specifically, rendering of all objects except the pointer is disabled, and the pointer's self-illumination effect is enhanced. This generates a rendered image, which is used as a mask for pointer segmentation.
[0074] Optical Character Recognition (OCR) technology includes text localization and text recognition. Similar to size detection, it generates bounding boxes for all text objects and associates the text content corresponding to the bounding boxes for recognition.
[0075] The proposed annotation framework comprehensively meets the needs of mainstream instrument reading algorithms. Through post-processing, specific requirements, such as pointer bounding boxes, can be extracted from existing annotations. Furthermore, this invention records fundamental instrument properties, including measurement range, the number of primary scale markers, and the interval between the pointer and the nearest primary scale marker. While these parameters are not essential for direct reading tasks, they contribute to building instrument-specific question-and-answer datasets, which are of significant value for developing multimodal foundational models in the instrumentation field.
[0076] Based on the aforementioned data synthesis and automatic annotation methods, this invention constructs the standardized benchmark SyncG. SyncG contains 16,000 training images and 4,000 test images. All images are rendered at a resolution of 1920×1080 using Blender's Cycles renderer with a sample size of 200. Through a GPU-accelerated pipeline, the generation, rendering, and annotation process for each image takes only 2-4 seconds, achievable using a single NVIDIA RTX 3090 GPU. SyncG includes 145 different environments, with 80% of the images employing varying camera perspectives. The measurement and appearance parameters of each instrument are randomly assigned, ensuring that different images can be considered to contain different instruments. Therefore, this invention successfully generates approximately 20,000 unique instruments.
[0077] Step 3: Scale marking and pointer prediction are performed using the DINOv2-based heatmap prediction model and key point extraction algorithm. In the key point extraction, a multi-stage extraction method is used for intermediate scale marking that may contain an uncertain number of key points.
[0078] This invention decomposes instrument reading into five distinct processing stages: instrument detection, scale and pointer detection, text detection and recognition, viewpoint correction, and final reading. The initial stage reduces environmental interference by separating the instrument from the complex background. Specifically, this invention uses YOLOv11 to locate the instrument. YOLOv11 is a lightweight, single-stage object detection network. This invention selects the minimum version of YOLOv11s and initializes the training process using officially pre-trained weights. Subsequently, the detected instrument region is cropped and resized to a standard size as input for subsequent processing stages. Therefore, before using the DINOv2-based heatmap prediction model and keypoint extraction algorithm for scale marking and pointer prediction, the process includes: locating the instrument using YOLOv11s and separating the instrument from the complex background.
[0079] Scale labeling and pointer prediction are performed using a DINOv2-based heatmap prediction model and keypoint extraction algorithm, specifically as follows:
[0080] In the scale marker and pointer prediction phase, the goal is to detect the positions of the pointer and scale markers within the meter. Five keypoint categories are selected, including: starting scale marker, ending scale marker, intermediate scale marker, the starting position of the pointer, and the ending position of the pointer. These scale marker keypoints are typically located near the center of the scale, while the pointer's starting point is located at the center of rotation, and its ending point is located far from the center. Traditional keypoint detection architectures are mainly used to solve a fixed number of prediction tasks, such as human pose estimation. However, due to the inherent variation in the number of main scale markers in different scale types, traditional methods are difficult to apply. Therefore, this invention develops a custom keypoint detection network and uses heatmap prediction as a proxy task. The training process first generates synthetic heatmaps, creating five different heatmap channels for each keypoint category. Next, this invention designs a heatmap prediction network designed to learn the mapping relationship from the original image to the heatmap. Finally, in the testing phase, this invention predicts the heatmap from the input image and extracts keypoints using a clustering algorithm.
[0081] Heatmap prediction models are used to predict heatmap values for each keypoint category. The heatmap prediction model is as follows:
[0082] This invention utilizes a frozen pre-trained DINOv2 encoder to extract features from the input image. A multi-stage upsampling network is employed to generate a heatmap of the same size as the original image, with BCE loss used as the prediction loss. Since the encoder reduces the input image by 1 / 14 in both height and width, this invention employs a multi-stage upsampling network instead of single-layer linear interpolation. Inspired by image super-resolution techniques, each upsampling stage consists of a convolutional layer and a transposed convolutional layer, the latter performing 2× upsampling. The final heatmap is generated through a single convolutional layer.
[0083] Keypoints are extracted from the heatmaps of each keypoint category. For intermediate scale labels, each heatmap corresponds to one keypoint, and the mean shift algorithm is used for extraction.
[0084] In keypoint extraction, a multi-stage extraction method is used for intermediate-scale markers that may contain an uncertain number of keypoints. Specifically:
[0085] Significant region filtering: Set a response threshold c, and retain only candidate regions whose heatmap responses satisfy H(x,y)≥c, thereby effectively suppressing background noise;
[0086] Local maximum extraction: A k×k sliding window is used to perform maximum filtering to detect local maxima that satisfy the following conditions:
[0087]
[0088] Non-maximum suppression: A greedy algorithm based on spatial distance constraints is adopted to iteratively select key points in descending order of confidence. Low-confidence candidate points within the radius r of the selected key points are eliminated to reduce the problem of key point overlap in dense areas.
[0089] Sub-pixel refinement: To improve positioning accuracy to the sub-pixel level, quadratic interpolation is used to adjust integer pixel coordinates, and peak offset is estimated using Taylor expansion.
[0090]
[0091] The keypoint coordinates (x, y) are optimized to (x+Δx, y+Δy), where both |Δx| and |Δy| do not exceed 0.5, thereby significantly improving accuracy at the sub-pixel level.
[0092] Step 4: After text detection and recognition, the homography matrix required for perspective correction is directly regressed from the original image end-to-end using a regression network. The relationship between the polar angle and the gauge reading is fitted using the RANSAC algorithm, and the value corresponding to the pointer polar angle is predicted through the fitted model.
[0093] In the text detection and recognition stage, the goal is to extract numerical text from the dashboard so that users can read the meter readings without knowing the measurement range. This task is divided into two sub-tasks: Text detection: This sub-task is responsible for locating and cropping the text region for further processing. Similar to the meter detection task, this invention uses YOLOv11s to achieve this goal. Text recognition: This sub-task extracts the text content from the cropped image. For this purpose, this invention uses ABINet from the MMOCR library.
[0094] Perspective correction is performed by directly regressing the homography matrix required for perspective correction from the original image end-to-end using a regression network. Specifically:
[0095] In practical applications, due to camera instability or space limitations, it is often impossible to make the camera perfectly perpendicular to the instrument panel. This deformation transforms the original circular dial into an ellipse, leading to errors in subsequent pointer angle calculations. To solve this problem, this invention employs perspective correction technology.
[0096] To avoid reliance on earlier stages, a deep learning model is trained using homography annotations in synthetic data to predict homography transformations.
[0097] From the original image to end-to-end processing. Since homography transformation can be represented by eight degrees of freedom, the model predicts a vector h∈R. 1×8The pre-trained DINO backbone network is used as the encoder to extract features from the input image and predict the final homography transformation through a linear layer, and is trained using the L1 loss function.
[0098] During inference, a 1 is added to the end of the prediction vector h, and it is reshaped into an isomorphic matrix suitable for images with normalized width and height. For an image of shape (h, w), the isomorphic transformation matrix H homo Calculate as follows:
[0099]
[0100] After obtaining the isomorphic transformation, the transformation is applied to the key points and text coordinates detected in the previous stages and mapped onto the corrected perspective plane for the final reading stage.
[0101] The process of obtaining the relationship between the polar angle and the gauge reading is as follows:
[0102] The inherent limitations of detection models, coupled with variations in image quality, can lead to inaccuracies in keypoint detection and OCR (including false positives and false negatives), rendering simple reading strategies unreliable. To address this issue, this invention proposes a more robust reading method. First, the OCR output is filtered based on confidence scores from both detection and recognition.
[0103] To more accurately determine the relative positions of the scale marker and the digital text, this invention establishes a polar coordinate system with the starting position of the pointer as the pole and the downward vertical direction as the polar axis, and calculates the polar angles of the scale marker, the pointer endpoint, and the center of the OCR bounding box.
[0104] Establish the relationship between the polar angle and the gauge reading, when the polar angle marked on the ruler... Polar angle with the center of the OCR bounding box The two will be matched if the following conditions are met:
[0105]
[0106] min(d ij ,(360-d ij ))≤d;
[0107] Where d is the threshold; when two or more matches are detected, the correspondence between the polar angles of these scale markers and their numerical values is recorded.
[0108] When fewer than two valid pairs are obtained, it indicates that there may be a failure in key point detection. The system will automatically switch back to the geometric cues derived from OCR. Based on experience, it has been observed that in a properly calibrated gauge, the angle of the scale mark is usually consistent with the angle of the OCR centroid.
[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0110] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying instrument readings based on a self-generated and annotated standard benchmark test set, characterized in that, include: Step 1: Use a Python script to randomly sample and set the measurement and appearance parameters, and use Blender to generate the instrument dataset; Step 2: Construct a fully automated specification annotation framework based on Blender 3D model parameters to annotate the instrument dataset and obtain the specification benchmark test set; Step 3: Scale marking and pointer prediction are performed using the DINOv2-based heatmap prediction model and key point extraction algorithm. In the key point extraction, a multi-stage extraction method is used for intermediate scale marking that may contain an uncertain number of key points. Step 4: After text detection and recognition, the homography matrix required for perspective correction is directly regressed from the original image end-to-end using a regression network. The relationship between the polar angle and the gauge reading is fitted using the RANSAC algorithm, and the value corresponding to the pointer polar angle is predicted through the fitted model.
2. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 1, the measurement parameters include: the rotation angle of the pointer and the measurement range of the instrument.
3. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 1, the appearance parameters include: the color and shape of the pointer, the length of the scale marks and the numerical spacing between adjacent marks, the material and corrosion condition of the instrument housing, the angle of the camera and its distance from the instrument, the style and size of the digital text and its radial position relative to the center, the texture and noise on the dial, and the diversity of the environment.
4. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, Step 2, the fully automated specification annotation framework based on Blender 3D model parameters, specifically includes: Gauge inspection and annotation of the outer bounding box: Traverse all vertices of the gauge outer shell in Blender and calculate the projected coordinates of the vertices in the camera image; convert the projected coordinates into a two-dimensional coordinate system with the top left corner as the origin, and summarize the maximum and minimum xy values of the projected coordinates to determine the bounding box of the gauge outer shell; Perspective correction: Establish the correspondence between two point sets: the positions of four cocircular equidistant points in the world coordinate system and the corresponding points of the camera projection of the four cocircular equidistant points; normalize both coordinate sets relative to the detected scale bounding boxes, and use OpenCV's isomorphic estimation method to calculate the perspective transformation matrix between the normalized coordinate systems, which serves as the training target for the distortion correction model. Scale mark detection: divided into key point detection and semantic segmentation; for key point detection, points adjacent to the center of the gauge are marked at each major scale mark; for semantic segmentation, polygon annotation is performed based on the geometry of the scale mark. Light spot detection is divided into key point detection and semantic segmentation. For key point detection, the intersection of the pointer rotation axis and the rotation axis surface is defined as the starting position, and the ending position is determined by traversing all points on the pointer and selecting the point farthest from the starting position. For pointer segmentation and annotation, a mask-based annotation method is used. Optical Character Recognition (OCR) technology includes text localization and text recognition. It generates bounding boxes for all text objects and associates the text content corresponding to the bounding boxes for recognition.
5. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, Step 3, before using the DINOv2-based heatmap prediction model and key point extraction algorithm for scale marking and pointer prediction, also includes: using YOLOv11s to locate the instrument and separating the instrument from the complex background.
6. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 3, the DINOv2-based heatmap prediction model and keypoint extraction algorithm are used for scale marking and pointer prediction, specifically: Five keypoint categories were selected, including: start scale marker, end scale marker, intermediate scale marker, pointer start position, and pointer end position; Heatmap prediction is performed for each keypoint category using a heatmap prediction model; the heatmap prediction model is as follows: The frozen pre-trained DINOv2 is used as an encoder to extract features from the input image; a multi-stage upsampling network is used to generate a heatmap of the same size as the original image, and BCE loss is used as the prediction loss; each upsampling stage consists of a convolutional layer and a transposed convolutional layer, the latter performing 2× upsampling, and the final heatmap is generated by a convolutional layer. Keypoints are extracted from the heatmaps of each keypoint category. For intermediate scale labels, each heatmap corresponds to one keypoint, and the mean shift algorithm is used for extraction.
7. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 3, a multi-stage extraction method is used for intermediate-scale markers that may contain an uncertain number of keypoints during keypoint extraction. Specifically: Significant region filtering: Set a response threshold c, and retain only candidate regions whose heatmap responses satisfy H(x,y)≥c; Local maximum extraction: A k×k sliding window is used to perform maximum filtering to detect local maxima that satisfy the following conditions: Non-maximum suppression: A greedy algorithm based on spatial distance constraints is adopted to iteratively select key points in descending order of confidence, and low-confidence candidate points within the radius r of the selected key points are eliminated; Subpixel thinning: Integer pixel coordinates are adjusted using quadratic interpolation, and peak offset is estimated using Taylor expansion. The keypoint coordinates (x, y) are optimized to (x+Δx, y+Δy), where both |Δx| and |Δy| do not exceed 0.
5.
8. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 4, YOLOv11s is used for text detection, and ABINet is used for text recognition.
9. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 4, a regression network is used to directly regress the homography matrix required for perspective correction from end to end in the original image to perform perspective correction. Specifically: A deep learning model is trained using homography annotations in synthetic data to predict homography transformations; Homography transformation is represented by eight degrees of freedom, and the model predicts a vector h∈R. 1×8 The pre-trained DINO backbone network is used as the encoder to extract features from the input image and predict the final homography transformation through a linear layer, and is trained using the L1 loss function. During inference, a 1 is added to the end of the prediction vector h, and it is reshaped into an isomorphic matrix suitable for images with normalized width and height. For an image of shape (h, w), the isomorphic transformation matrix H homo Calculate as follows: After obtaining the isomorphic transformation, it is applied to the key points and text coordinates detected in the previous stages and mapped onto the corrected perspective plane.
10. The instrument reading identification method based on a self-generated and annotated standard benchmark test set according to claim 1, characterized in that, In step 4, the process of obtaining the relationship between the polar angle and the gauge reading is as follows: The OCR output is filtered based on confidence scores by combining detection and recognition. A polar coordinate system was established with the starting position of the pointer as the pole and the downward vertical direction as the polar axis. The polar angles of the scale mark, the pointer endpoint, and the center of the OCR bounding box were calculated. Establish the relationship between the polar angle and the gauge reading, when the polar angle marked on the ruler... Polar angle with the center of the OCR bounding box The two will be matched if the following conditions are met: min(d ij ,(360-d ij ))≤d; Where d is the threshold; when two or more matches are detected, the correspondence between the polar angles of these scale markers and their numerical values is recorded.
Citation Information
Patent Citations
Visual tracking and annotation of clinically important anatomical landmarks for surgical interventions
CN102781336A
Method, apparatus, and system for initializing a meter reading device
CN107076571A
A neural network for head pose and gaze estimation using photorealistic synthetic data
CN114041175A
System, method and device for reading measuring device, and storage medium
CN115708133A
Heat supply data visualization method based on digital twinning and computer equipment
CN117742858A