Dial reading recognition method
Patent Information
- Application Number
- CN202611232049.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,在实际变电站巡检中,巡检机器人位姿变化、手持设备角度偏差不可避免,导致采集到的表盘图像存在显著的透视变形
Smart Images

Figure CN122821180A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power operation and maintenance technology, and in particular to a method for recognizing meter readings. Background Technology
[0002] Substations are equipped with a large number of pointer-type instruments (voltmeters, ammeters, oil temperature gauges, etc.), which are the core monitoring carriers for reflecting the operating conditions of equipment and ensuring the safety of the power grid. Currently, related technologies, through classification models and key point detection, or through template matching and edge detection, can achieve a certain degree of reading recognition in standard frontal or small-angle deviation scenarios.
[0003] However, in actual substation inspections, changes in the robot's pose and deviations in the angle of the handheld device are unavoidable, resulting in significant perspective distortion in the acquired dial images. Since the relevant technologies rely on distorted images for pointer extraction, center location, and angle calculation, a fixed, systematic deviation arises between the pointer deflection angle and the true value. Furthermore, the more severe the distortion, the greater the reading error, thus limiting the accuracy of dial reading recognition. Summary of the Invention
[0004] Therefore, it is necessary to provide a dial reading recognition method that can reduce systematic errors caused by perspective distortion, in order to address the above-mentioned technical problems.
[0005] This application provides a method for recognizing dial readings, including:
[0006] Acquire the dial image of the watch face to be read, and extract the feature vector of the dial image;
[0007] The dial image and the target template image are input into a feature extraction network to obtain multi-scale feature maps of the dial image and the target template image. For each location in the multi-scale feature maps of the dial image and the target template image, feature information from other locations is fused to obtain a first global feature of the dial image and a second global feature of the target template image. The multi-scale feature maps include at least a low-resolution feature map and a high-resolution feature map. The low-resolution feature map is used to represent the global structural information and global semantic information of the corresponding image, while the high-resolution feature map is used to represent the local structural information and local semantic information of the corresponding image.
[0008] Cross-attention matching is performed on the first global feature and the second global feature to obtain multiple sets of matching point pairs;
[0009] The feature vector is matched with each template feature vector in the preset template feature library to find the target template image that matches the dial image; the template feature library stores at least one standard dial template image, a template feature vector corresponding to each standard dial template image, and a scale reference parameter corresponding to each standard dial template image;
[0010] The dial image is matched with the target template image by feature point matching to obtain multiple sets of matching point pairs;
[0011] Based on the multiple sets of matching point pairs, the perspective transformation parameters of the dial image relative to the target template image are determined;
[0012] According to the perspective transformation parameters, the dial image is subjected to inverse perspective transformation to obtain the transformed dial image;
[0013] The pointer position is determined in the transformed dial image, and the dial reading of the dial to be read is determined according to the pointer position and the scale reference parameters corresponding to the target template image.
[0014] The aforementioned dial reading recognition method first determines the target template image by matching the feature vector of the dial to be tested with a template feature library. Based on this, feature point matching is performed between the dial to be tested and the target template image, resulting in multiple sets of matching point pairs. This establishes a pixel-level geometric correspondence between the dial image of the dial to be read and the standard dial template image, allowing for accurate quantification of the degree and shape of distortion in the dial image of the dial to be read. Subsequently, perspective transformation parameters are determined based on the multiple sets of matching point pairs. Then, an inverse perspective transformation is performed on the distorted image according to the perspective transformation parameters. This application performs an inverse perspective transformation on the dial image of the dial to be read to obtain a transformed dial image, reducing perspective distortion in the image space. Finally, the pointer position is determined in the transformed dial image, and the reading is read, improving the accuracy of dial reading recognition. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is an application environment diagram of a dial reading recognition method in one embodiment;
[0017] Figure 2 This is a flowchart illustrating a dial reading recognition method in one embodiment;
[0018] Figure 3 This is a schematic diagram illustrating the process of determining multiple sets of matching point pairs in one embodiment;
[0019] Figure 4 This is a schematic diagram of the complete process of a dial reading recognition method in another embodiment;
[0020] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0022] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0023] The dial reading recognition method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 can acquire an image of the dial to be read and send it to server 104. After receiving the dial image, server 104 extracts the feature vector of the dial image and matches it with the feature vectors of each template in the template feature library to find the target template image. Then, server 104 performs feature point matching between the dial image and the target template image, obtaining multiple sets of matching point pairs, and determines the perspective transformation parameters accordingly. It then performs an inverse perspective transformation on the dial image to obtain the transformed dial image. Finally, server 104 determines the pointer position in the transformed dial image and, combined with the scale reference parameters corresponding to the target template image, calculates the dial reading and returns the dial reading to terminal 102. Alternatively, terminal 102 can also perform all the above steps locally without relying on server 104. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection equipment. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0024] In one exemplary embodiment, such as Figure 2 As shown, a dial reading recognition method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 201 to 207. Wherein:
[0025] Step 201: Obtain the dial image of the dial to be read and extract the feature vector of the dial image.
[0026] The dial image refers to the original image of the dial area of the pointer-type meter to be read, captured by the meter terminal through a camera.
[0027] First, the server receives raw inspection images from substation inspection front-end equipment (such as industrial cameras mounted on inspection robots or fixed-point monitoring cameras), and calls a pre-trained object detection model to perform dial target detection on the raw inspection images. A sub-image containing only the effective area of the dial to be read is cropped from the raw inspection images as the dial image. Simultaneously, the dial image is uniformly scaled to a preset size to adapt to the input requirements of the subsequent feature extraction network. In this embodiment, the object detection model uses a lightweight YOLO detection model with MobileNetV3 as the backbone network. The input of the object detection model is the raw substation inspection image, and the output is the coordinates of the dial target bounding box and its corresponding confidence score. During the training phase, the object detection model uses multiple dial images collected and labeled in various substation scenarios as the training dataset, covering common meter types such as voltmeters, ammeters, oil temperature gauges, and pressure gauges, and is divided into training, validation, and test sets according to proportions. For example, during training, the CIoU (Complete Intersection over Union) loss function is used as the bounding box regression loss, combined with Focal Loss as the classification loss, and trained for multiple epochs according to a preset initial learning rate. After training, the model weights that perform best on the validation set are selected as the pre-trained object detection model.
[0028] Then, the server calls the pre-trained feature extraction model, inputting the dial image of the meter to be read into the feature extraction model for forward inference calculation. The feature extraction model extracts high-dimensional global feature representations from the dial image through multi-layer convolution and pooling operations, and finally outputs a fixed-dimensional feature vector. The feature vector is used to represent the appearance and texture features of the dial. The feature extraction model has the exact same network structure and training parameters as the feature extraction model used in the subsequent template feature library construction, thus ensuring that all feature vectors are in a unified high-dimensional embedding space, making the cosine similarity comparison have a reasonable one-to-one correspondence. In this embodiment, the feature extraction model adopts the Fast-ReID architecture based on the ResNet-50-ibn backbone network. The input of the feature extraction model is a cropped and size-normalized dial image, and the output is a 2048-dimensional feature vector. During the training phase, the feature extraction model uses a dataset of substation dial images containing various models, ranges, lighting conditions, and shooting angles, divided into training, validation, and test sets according to proportions. During training, a joint optimization strategy of Circle Loss and Center Loss is employed. Specifically, for the Fast-ReID feature extraction model, the joint optimization of Circle Loss and Center Loss is proposed to address the problem of poor adaptability of general ReID models to dial features, that is... In the formula, This represents the total value of the joint loss function. The value of the Circle Loss function; The value of the Center Loss function; These are the weighting coefficients for Center Loss, used to balance the contributions of the two types of loss to the total loss.
[0029]
[0030] , The numerical value of the Center Loss function; This refers to the training batch size; For batch sample index For the first Feature vectors of each sample; For the first The category to which each sample belongs Corresponding category feature centers; For the first The category labels of each sample; It is the square of the Euclidean norm.
[0031] As can be seen from the two equations above, Circle Loss focuses on optimizing the relative ranking relationship in the feature embedding space, that is, bringing positive sample pairs closer while pushing negative sample pairs further apart. However, its constraint on the class feature center is implicit and indirect. The introduction of Center Loss makes up for this deficiency. Center Loss maintains a class feature center for each meter category. And during training, the feature vectors of all samples of the same class are forced to... To its corresponding category center To move closer. For example, the two... The weight combination forms a dual regularization effect that combines macro-level ranking optimization and micro-level clustering constraints. Circle Loss ensures that different categories are globally separable in the embedding space, while Center Loss ensures that the same category is locally compact in the embedding space. Since the appearance of meters of the same model varies under different lighting and angles, and the dials of different models and ranges may look extremely similar, the joint loss function improves the model's feature discrimination ability in difficult scenarios with large intra-class differences and small inter-class differences.
[0032] The optimizer used is Adam (Adaptive Moment Estimation), with a preset initial learning rate, which decays to 0.1 times the current rate at the 50th and 100th epochs, for a total of the preset number of training epochs. After training, the model weights that perform best on the validation set are selected as the trained feature extraction model.
[0033] Step 202: Match the feature vector with each template feature vector in the preset template feature library to find the target template image that matches the dial image; the template feature library stores at least one standard dial template image, the template feature vector corresponding to each standard dial template image, and the scale reference parameters corresponding to each standard dial template image.
[0034] The scale reference parameters include the dial center position, the starting pixel position of the scale, the ending pixel position of the scale, the lower limit of the range and the upper limit of the range, as well as geometric parameters such as the inner and outer diameters of the effective scale area, which are used to completely describe the scale space distribution of the dial.
[0035] The server employs a cosine similarity algorithm to calculate the cosine of the angle between the feature vector of the dial image and the template feature vector of each standard dial template image in the template feature library. This cosine value is used as a quantitative score to measure the degree of feature similarity between the two images. Specifically, the feature vector of the dial image and the template feature vector of each standard dial template image in the template feature library are normalized by their magnitudes. Then, the inner product of the two normalized vectors is calculated, and the result is the similarity score, ranging from -1 to 1. The closer the value is to 1, the more consistent the directions of the two feature vectors are in the embedding space, meaning the more similar the appearance features of the two dial images are. The server iterates through all standard dial template images in the template feature library, calculates a similarity score for each standard dial template image, and selects the template with the highest score as a candidate matching point pair template. Simultaneously, the server compares the highest similarity score with a preset matching threshold. When the highest similarity score is lower than the matching threshold, it determines that the meter type corresponding to the current dial image is not registered in the template feature library. At this time, the server generates an unknown meter alarm and terminates the subsequent recognition process to avoid incorrect matching. When the highest similarity score reaches or exceeds the matching threshold, the standard dial template image corresponding to the highest similarity score is determined as the target template image, and the scale reference parameters associated with the target template image are read from the template feature library for use in subsequent steps. It should be noted that the above-mentioned matching threshold values and judgment methods are only examples. In practical applications, they can be adaptively adjusted according to factors such as meter type, scene lighting conditions, and recognition accuracy requirements. This application is not limited to this.
[0036] Step 203: Input the dial image and the target template image into a feature extraction network to obtain multi-scale feature maps of the dial image and the target template image. For the feature information of each position in the multi-scale feature maps of the dial image and the target template image, fuse the feature information of other positions except the stated position to obtain the first global feature of the dial image and the second global feature of the target template image. The multi-scale feature map includes at least a low-resolution feature map and a high-resolution feature map. The low-resolution feature map is used to represent the global structural information and global semantic information of the corresponding image, and the high-resolution feature map is used to represent the local structural information and local semantic information of the corresponding image.
[0037] The server simultaneously inputs the dial image and the target template image into a pre-trained dense feature matching network. This network extracts multi-scale depth features from both images using a deep neural network and establishes dense associations between their features through a global attention mechanism, ultimately outputting multiple reliable matching point pairs. Specifically, the server inputs the dial image and the target template image into a pre-trained feature extraction network. This network contains multi-level convolutional structures, enabling it to downsample the input image at each level, outputting feature maps of different spatial resolutions at each level. These feature maps of different resolutions collectively constitute a multi-scale feature map. The low-resolution feature map, due to multiple downsampling operations, has a large receptive field, with each pixel incorporating information from a large area of the image. Therefore, it excels at capturing global structural information such as the dial's circular outline and overall layout, as well as abstract global semantic information about the dial type. The high-resolution feature map, on the other hand, retains more spatial details and has a smaller receptive field, accurately responding to fine local structural and semantic information such as the edges of the scale lines, the tips of the hands, and the strokes of the numbers. This allows subsequent global feature fusion and cross-attention matching to be performed in a feature space that simultaneously contains both "what the whole is" and "where the details are," achieving registration that is coarse-to-fine and considers both semantics and geometry. Therefore, feature maps at multiple scales can utilize global contextual information while also taking into account local fine features.
[0038] The matching process described above is achieved by using a neural network to learn the pixel-level semantic correspondence between two images. Even when the dial has changes in lighting, partial occlusion, or slight rotation, it can still output a sufficient number of spatially evenly distributed matching point pairs. The server performs preliminary quality screening on the output matching point pairs, discarding matches with low confidence and retaining high-reliability matching point pairs for subsequent transformation parameter estimation.
[0039] It should be noted that the dense feature matching network used in this embodiment is a Transformer-based LoFTR network. The dense feature matching network consists of two parts: a feature extraction network and a matching module. The feature extraction network uses a shared-weight CNN (Convolutional Neural Network) backbone (ResNet-101) to extract multi-scale feature maps from the two input images, outputting a feature map with a resolution equal to the input image. The matching module consists of multiple stacked self-attention layers and bidirectional cross-attention layers, used to establish pixel-level feature correspondences between the two images and output a matching score matrix. The LoFTR network takes two watch face images to be matched as input and outputs multiple sets of matching point pairs. Each set of matching point pairs includes the coordinates of feature points in the watch face image and their corresponding matching point coordinates in the target template image, as well as the corresponding matching confidence score. The dense feature matching network is pre-trained on the MegaDepth outdoor scene dataset, and a loss function is used to supervise the matching probability during the pre-training stage. Subsequently, the constructed dial dataset was fine-tuned using paired data containing multiple substation dial images and template images, with the same loss function applied during the fine-tuning phase. After fine-tuning, the model weights that performed best on the validation set were selected as the dense feature matching network.
[0040] The server performs self-attention processing on each spatial location in the multi-scale feature map of the dial image and each spatial location in the multi-scale feature map of the target template image. Specifically, for the multi-scale feature map of the dial image, the server treats each spatial location as a target location and calculates the feature similarity between the target location and all other locations in the same feature map to obtain a set of attention weights. This set of attention weights reflects the correlation strength between the target location and other global locations. Subsequently, based on this set of attention weights, the server weighted and aggregated the feature information of all other locations and fused the aggregated global context information into the original feature expression of the target location, thereby obtaining a new feature of the target location after global enhancement. The above operation is performed on each spatial location in the multi-scale feature map of the dial image to obtain the first global feature. It can be understood that in the first global feature, the feature of each location is fused with global structural and semantic information from the entire image. Similarly, the same self-attention processing is performed on the multi-scale feature map of the target template image to obtain the second global feature, in which each location is fused with global information from the entire template image. That is to say, each location in the first and second global features has the ability to perceive the entire image globally.
[0041] Step 204: Perform cross-attention matching on the first global feature and the second global feature to obtain multiple sets of matching point pairs.
[0042] In this context, a matching point pair refers to a one-to-one correspondence between a pixel in the dial image and its corresponding pixel in the target template image, representing the visual correspondence between the two images at the same spatial location. Cross-attention matching models the image matching process as a bidirectional interaction and information aggregation process between feature sequences.
[0043] The server takes a first global feature and a second global feature as input. By calculating the similarity between any positions in the first and second global features, it generates a matching score matrix that describes the strength of the association between features, thus explicitly quantifying the degree of matching between each position in the test image and each position in the template image. Based on the matching score matrix, it can filter and confirm matching point pairs in a learnable or optimizable manner by integrating the contextual information of the entire dial image. This ensures that each matching decision depends not only on the feature similarity between the two positions to be matched but also on the global constraints of the matching relationships of other positions. When there are many repetitive textures (such as similarly arranged tick marks) or local occlusions in the dial image, it can mitigate erroneous matches caused by the lack of global awareness in local descriptors.
[0044] Step 205: Based on multiple sets of matching point pairs, determine the perspective transformation parameters of the dial image relative to the target template image.
[0045] Among them, the perspective transformation parameter is used to describe the homography transformation relationship between the plane where the dial image is located and the plane where the target template image is located. The homography transformation relationship characterizes the perspective distortion of the dial caused by the tilt of the shooting angle.
[0046] The server employs a random sampling consensus algorithm, repeatedly and randomly selecting several pairs of points from multiple sets of matched point pairs to estimate candidate perspective transformation matrices. All matched point pairs are then substituted into the candidate transformation matrices to calculate the projection error. The number of point pairs with errors less than a preset threshold is counted as the number of inliers. After multiple iterations, the server selects the candidate transformation matrix with the highest number of inliers as the final perspective transformation parameter. The projection error of the m-th matched point pair is... The calculation can be shown in the following formula:
[0047]
[0048] In the formula, For the target template image, the first Cartesian coordinates of the matching points; A function for converting homogeneous coordinates to Cartesian coordinates; for ; The first one in the image of the dial to be tested Homogeneous coordinates of matching points; It is the Euclidean norm.
[0049] From the above formula, we can see that the projection error The physical meaning is in the currently estimated perspective transformation model The following describes the spatial offset between the theoretical projection position and the actual matching position after projecting the coordinates of the matching point in the dial under test onto the target template image space. The projection error measures the impact of the current estimated transformation model on the first... Goodness of fit of matched point pairs.
[0050] Step 206: Perform inverse perspective transformation on the dial image according to the perspective transformation parameters to obtain the transformed dial image.
[0051] Inverse perspective transformation refers to remapping the pixels of a dial image according to the inverse transformation relationship of perspective transformation parameters, thereby correcting the dial image that was originally distorted by perspective due to the tilt of the shooting angle into a dial image with a standard frontal view.
[0052] The server constructs a blank output image with the same size as the target template image. For each pixel position in the output image, it calculates the corresponding sub-pixel coordinates in the original dial image based on the inverse matrix of the perspective transformation parameters. Then, it samples pixel values from the original dial image and fills them into the corresponding positions in the output image using bilinear interpolation to generate the transformed dial image.
[0053] It is understood that the transformed dial image described here is the transformed dial image obtained after perspective distortion correction of the dial to be read. After inverse perspective transformation, the dial plane and the imaging plane are restored to a parallel relationship, and the scale distribution and pointer direction on the dial are restored to their true geometric shape under standard frontal viewing conditions. This provides a distortion-free image basis for the subsequent accurate extraction of pointer position and reading calculation. In other words, the transformed dial image is equivalent to a transformed dial image of the dial to be read.
[0054] Step 207: Determine the pointer position in the transformed dial image, and determine the dial reading of the dial to be read based on the pointer position and the scale reference parameters corresponding to the target template image.
[0055] The pointer position refers to the spatial orientation of the dial pointer in the transformed dial image, including the position of the pointer tip and the direction the pointer points.
[0056] The server invokes a pre-trained instance segmentation model to perform forward inference on the transformed dial image. The instance segmentation model outputs a binary mask corresponding to the pointer region, identifying the set of pixels in the image belonging to the pointer. In this embodiment, the instance segmentation model adopts the YOLOv11-seg architecture. The input is the dial image after inverse perspective transformation, and the output is a binary mask of the pointer region (with the same resolution as the input image) and the corresponding class confidence score. During the model training phase, multiple dial images collected and labeled in various substation scenarios are used as the training dataset. Each image is labeled with a pixel-level outline polygon of the pointer region, and the dataset is divided into training, validation, and test sets according to proportions. During training, a joint loss function of mask prediction loss, bounding box regression loss CIoU, and classification loss Focal Loss is used for end-to-end training. A total of preset training epochs are performed. After training, the model weights with the best performance on the validation set are selected as the pre-trained instance segmentation model. Understandably, the pointer region provided by the mask is a coarse-grained representation at the pixel level, while the pointer position required for the final reading calculation is fine-grained geometric pointing information. Therefore, it is necessary to further extract the pointer's geometric parameters from the pointer region. Based on this, the server extracts the pointer's pointing direction according to the geometry of the pointer region. In one example, the minimum bounding rectangle of the pointer region is calculated based on the mask, or the skeleton is refined to obtain the pointer tip coordinates. Combined with the dial center coordinates marked in the target template image, the deflection angle from the center to the pointer tip is calculated. The deflection angle is also the quantized expression of the pointer position in angle space. Then, the server maps the starting and ending pixel positions of the scale in the target template image to the transformed dial image coordinate space. Combining the proportional relationship between the deflection angle and the angle interval between the starting and ending scales, as well as the lower and upper limits of the range in the scale reference parameters, the dial reading of the dial to be read is calculated through linear interpolation.
[0057] In this embodiment, the target template image is first determined by matching the feature vector of the dial to be tested with a template feature library. Then, feature point matching is performed between the dial to be tested and the target template image to obtain multiple sets of matching point pairs. A pixel-level geometric correspondence is established between the dial image of the dial to be read and the standard dial template image, allowing the degree and shape of distortion in the dial image of the dial to be read to be accurately quantified. Subsequently, perspective transformation parameters are determined based on the multiple sets of matching point pairs. Then, an inverse perspective transformation is performed on the distorted image according to the perspective transformation parameters. This application performs an inverse perspective transformation on the dial image of the dial to be read to obtain a transformed dial image, reducing perspective distortion in the image space. Finally, the pointer position is determined in the transformed dial image and the reading is read, improving the recognition accuracy of the dial reading.
[0058] Figure 3For a schematic diagram of the process of determining multiple sets of matching point pairs in one embodiment, please refer to [link / reference]. Figure 3 In one exemplary embodiment, feature point matching is performed between the dial image and the target template image to obtain multiple sets of matching point pairs, including:
[0059] Step 301: Perform cross-attention processing on the first global feature and the second global feature to establish the feature correspondence between the dial image and the target template image.
[0060] First, the server uses the first global feature of the dial image as the query feature and the second global feature of the target template image as the key feature. By calculating the relevance score between the query feature and the key feature, the server obtains the relevance score distribution between each position in the dial image and all positions in the target template image. The relevance score distribution quantifies the semantic similarity between each position in the dial image and all positions in the target template image.
[0061] Then, the server uses the relevance score as an attention weight, applying it to the second global feature of the target template image. Through weighted aggregation, it obtains cross-image association features corresponding to each position in the dial image. Simultaneously, the server performs the reverse process: using the second global feature of the target template image as the query feature and the first global feature of the dial image as the key-value feature, it calculates the relevance score distribution between each position in the target template image and each position in the dial image, and performs reverse weighted aggregation accordingly, achieving bidirectional cross-image feature interaction. After this cross-attention processing, the features at each position in both the dial and target template images contain information about related positions in the other image, thus establishing a dense feature correspondence between the two images globally.
[0062] It is understandable that the feature correspondence established by cross-attention processing is an abstract concept at the logical level, representing the semantic association and matching probability between pixel positions in two images; while the correlation score distribution is the specific numerical representation of the feature correspondence, and the level of the correlation score is used to quantify the strength of the feature correspondence. The feature correspondence focuses on describing the fact that "which positions correspond to each other" between two images, while the correlation score distribution specifically expresses, in the form of a numerical matrix, "how high the confidence level is between each pair of positions".
[0063] Step 302: Based on the feature correspondence, select mutually matching feature points from the dial image and the target template image to obtain multiple sets of matching point pairs.
[0064] The server selects the position with the highest correlation score in the target template image for each position in the dial image as a candidate matching point pair based on the correlation score between each position in the target template image and all positions in the target template image. Simultaneously, the server performs a reverse consistency check on the candidate matching point pairs based on the correlation score between each position in the target template image and all positions in the dial image. This checks whether the correlation score between the candidate position and its original position in the dial image is also the highest. Only position pairs that pass the bidirectional consistency check are retained as valid matches. The server sorts all candidate matching point pairs that pass the check from highest to lowest correlation score and selects several pairs with scores higher than a preset threshold as the final matching point pairs output. The matching point pairs are evenly distributed and have high confidence in both the dial image and the target template image, accurately reflecting the geometric correspondence between the pixel positions in the two images.
[0065] In this embodiment, a self-attention mechanism enables each location to acquire global perception capabilities, and a bidirectional cross-attention mechanism is used to automatically learn the pixel-level correspondence between the global context of two images. This allows for the stable output of sufficient and highly reliable matching point pairs even under complex conditions such as changes in lighting, partial occlusion, slight rotation, and scale differences on the dial.
[0066] In an exemplary embodiment, based on feature correspondence, matching feature points are selected from the dial image and the target template image to obtain multiple sets of matching point pairs. This includes: constructing a matching score matrix based on feature correspondence; the element values in the matching score matrix are used to characterize the degree of matching between feature points in the dial image and feature points in the target template image; optimizing the matching score matrix based on the element values in the matching score matrix to obtain an optimized matching score matrix; the optimized matching score matrix is used to establish a globally optimal one-to-one correspondence between feature points in the dial image and feature points in the target template image; and selecting matching feature points from the dial image and the target template image according to the optimized matching score matrix to obtain multiple sets of matching point pairs.
[0067] The matching score matrix is a two-dimensional matrix. The row index of the matching score matrix corresponds to the position of each feature point in the dial image, and the column index corresponds to the position of each feature point in the target template image. Each element value in the matrix is used to characterize the degree of matching between the feature point at the corresponding position in the dial image and the feature point at the corresponding position in the target template image.
[0068] First, the server fills the matching score matrix with the correlation scores between each position in the dial image and all positions in the target template image, according to the row and column correspondence. The higher the correlation score, the more likely the pair of feature points are to form a correct match semantically and spatially. Simultaneously, the server uses the correlation scores obtained from the reverse cross-attention processing as a consistency constraint to perform bidirectional verification on the elements in the matching score matrix. That is, only when both positive and negative correlations are high will the corresponding element value be assigned a high matching score, thus enabling the matching score matrix to comprehensively reflect the bidirectional feature correspondence.
[0069] Then, the server optimizes the matching score matrix based on the values of each element, obtaining an optimized matching score matrix. The goal of the optimization is to establish a globally optimal one-to-one correspondence between feature points in the dial image and feature points in the target template image. This ensures that each feature point in the dial image corresponds to at most one feature point in the target template image, and vice versa, with the total matching score of all matching pairs reaching the global maximum. In practice, the server models the matching point pair selection problem as solving a globally optimal transmission matching problem based on optimal transmission theory. Specifically, it seeks the optimal double random matching matrix that optimizes the overall matching score between feature points in the dial image and feature points in the target template image. The server uses the values of each element in the matching score matrix as the initial matching cost and iteratively normalizes the matrix by alternating rows and columns using the Sinkhorn algorithm. The iterative process continuously adjusts the values of each element in the matrix, gradually ensuring that the matching score matrix satisfies the normalization constraint in both the row and column directions. This means the sum of elements in each row and column of the matching score matrix is 1, resulting in a double-random matching matrix. This double-random matching matrix is the optimized matching score matrix. The optimized matching score matrix maximizes the global matching score while satisfying the one-to-one row-to-column correspondence constraint. It also preserves high-confidence matching relationships from the original matching score matrix, eliminates conflicting matches, and establishes a globally optimal bidirectional one-to-one correspondence between feature points in the dial image and feature points in the target template image.
[0070] Finally, the server iterates through the optimized matching score matrix. For each non-zero element in the optimized matching score matrix, it forms a matching point pair by pairing the feature point position in the dial image corresponding to that element's row with the feature point position in the target template image corresponding to that column. Simultaneously, the server selects matching point pairs whose element values are higher than a preset threshold as the final valid matching point pairs for output, based on the magnitude of the element values in the optimized matrix. These matching point pairs have global consistency in spatial distribution, exhibit no cross-conflicts, and accurately and uniquely reflect the geometric correspondence of pixel positions between the two images.
[0071] In this embodiment, based on the bidirectional cross-attention feature correspondence, the server can effectively eliminate matching ambiguity by constructing a matching score matrix and performing global optimization, ensuring that the selected matching point pairs satisfy a one-to-one correspondence in the global scope, thereby improving the accuracy and reliability of matching.
[0072] In an exemplary embodiment, determining the perspective transformation parameters of a dial image relative to a target template image based on multiple sets of matching point pairs includes: selecting a set of matching point pairs from the multiple sets of matching point pairs; determining initial perspective transformation parameters based on the selected matching point pairs; projecting each matching point in the dial image onto the coordinate space of the target template image according to the initial perspective transformation parameters to obtain the projection position corresponding to each matching point; calculating the positional error between the projection position corresponding to each matching point and the corresponding matching point in the target template image; determining matching point pairs with positional errors less than a preset error threshold as valid matching point pairs; updating the initial perspective transformation parameters according to the number of valid matching point pairs; and returning to the step of selecting a set of matching point pairs from the multiple sets of matching point pairs until a preset iteration stop condition is met to obtain the perspective transformation parameters.
[0073] Before formally beginning the iterative solution, it is necessary to understand the overall technical strategy adopted in this embodiment. The aforementioned embodiments obtained a sufficient number of spatially evenly distributed sets of matching point pairs, constituting a high-quality dense set of matching point pairs. Based on this, this embodiment employs a robust estimation algorithm based on RANSAC (Random Sample Consensus) and combines it with the DLT (Direct Linear Transform) method to solve for the homography matrix. Specifically, the RANSAC algorithm is responsible for separating correct matches (i.e., interior points conforming to the true perspective transformation relationship) from incorrect matches (i.e., mismatched exterior points caused by local similarity) in the matching point pair set through iterative random sampling and interior point screening mechanisms, thereby eliminating the interference of exterior points on parameter estimation. The DLT method is responsible for solving for candidate homography matrices in each iteration based on the currently selected matching point pairs using a system of linear equations. The two work together to ultimately estimate the optimal homography matrix, i.e., the perspective transformation parameters, from the dense set of matching point pairs.
[0074] First, the server randomly selects a preset number of matching point pairs from all matching point pairs as a sample subset, and the matching point pairs are spatially dispersed as much as possible to avoid collinearity and improve the stability of subsequent parameter estimation.
[0075] Then, based on the selected matching point pairs, the server uses the DLT method to solve for the initial perspective transformation parameters. Specifically, the mathematical relationship of perspective transformation is a homogeneous coordinate mapping relationship from a two-dimensional plane to a two-dimensional plane, expressed in the form of a homography matrix. The homography matrix contains multiple degrees of freedom, and its function is to map any pixel coordinates on the dial image plane to the corresponding coordinates on the target template image plane. The mapping relationship can uniformly describe various geometric transformations such as translation, rotation, scaling, tilting, and perspective distortion. During the solution process, the server represents the coordinates of the dial image matching points and the target template image matching points in each selected matching point pair as homogeneous coordinates, and uses the mapping correspondence between the coordinates of each pair of matching points to establish linear equations about the elements of the homography matrix. Each pair of matching points can contribute two linear equations. When the number of selected matching point pairs reaches the minimum number required for the solution, the server combines all the linear equations corresponding to multiple pairs of matching points into a homogeneous linear equation system. Because perspective transformation relationships have scale uncertainty, the elements of the homography matrix differ by at most one non-zero constant factor. Therefore, the server applies normalization constraints to the perspective transformation parameters to eliminate scale ambiguity. The server obtains the numerical solutions of each element in the homography matrix by solving a system of homogeneous linear equations. These solutions are the initial perspective transformation parameters, used to describe the homography relationship between the dial image plane and the target template image plane.
[0076] Next, the server uses the original coordinates of each matching point in the dial image as input, applies the homography matrix corresponding to the initial perspective transformation parameters to the coordinates of each matching point, and maps the coordinates of the matching points in the dial image to the coordinate space of the target template image through matrix multiplication, obtaining the projection position of each matching point in the coordinate space of the target template image. The projection position represents the theoretical spatial position that each matching point in the dial image should correspond to in the target template image under the current initial perspective transformation parameters.
[0077] Then, the server calculates the positional error between the projected position of each matching point and the corresponding point in the target template image that actually matches that matching point. Specifically, for each matching point in the dial image, the server compares its theoretical projected coordinates obtained after projecting it using the initial perspective transformation parameters with the coordinates of the actual matching point in the target template image corresponding to that matching point, and calculates the Euclidean distance between the two. The Euclidean distance is the positional error corresponding to that matching point. The magnitude of the positional error reflects the degree of fit between this set of matching point pairs and the overall transformation model under the currently estimated perspective transformation parameters.
[0078] Next, the server identifies matching point pairs with positional errors less than a preset error threshold as valid matching point pairs. When the positional error of a matching point pair is less than this threshold, it indicates that the geometric relationship of the matching point pair is highly consistent with the currently estimated perspective transformation parameters, and the matching point pair belongs to the interior points that conform to the overall transformation model; conversely, matching point pairs with positional errors greater than or equal to the threshold are considered exterior points and are discarded. Through the above filtering, the server effectively separates interior points from exterior points in the set of matching point pairs, thereby eliminating the interference of erroneous matching point pairs in the transformation parameter estimation.
[0079] Then, the server updates the initial perspective transformation parameters based on the number of valid matching point pairs. When the number of valid matching point pairs reaches a preset requirement, the server recalculates the perspective transformation parameters using all valid matching point pairs. That is, using all valid matching point pairs as the input set, the DLT method is used again to establish a system of linear equations and solve the transformation matrix to obtain updated, more accurate perspective transformation parameters. When the number of valid matching point pairs does not reach the preset requirement, the server discards the current initial perspective transformation parameters, does not update them, and directly proceeds to the next iteration.
[0080] Finally, the server returns to the step of selecting a set of matching point pairs from multiple sets of matching point pairs, repeating the above process of random selection, initial transformation parameter estimation based on DLT, projection calculation, position error comparison, interior point screening, and parameter update based on all interior points, until a preset iteration stopping condition is met. In this embodiment, the preset iteration stopping condition may include the following two cases: first, the number of iterations reaches a preset maximum iteration limit; second, in a certain round of iteration, the proportion of effective matching point pairs in all matching point pairs exceeds a preset proportion threshold, indicating that the currently estimated perspective transformation parameters have sufficient confidence. When either of the above stopping conditions is met, the server terminates the iteration and outputs the currently retained optimal perspective transformation parameters as the perspective transformation parameters. It should be noted that the above examples of iteration stopping conditions are given only for ease of understanding and do not constitute the only limitation of this application.
[0081] In this embodiment, the iterative sampling and interior point screening process effectively eliminates mismatched exterior points in the set of matching point pairs, thereby eliminating the interference of incorrect matching on the estimation of perspective transformation parameters and ensuring that the final solved perspective transformation parameters have globally optimal geometric consistency.
[0082] In an exemplary embodiment, the dial image is subjected to inverse perspective transformation according to perspective transformation parameters to obtain a transformed dial image, including: obtaining inverse transformation parameters of the perspective transformation parameters; determining the mapping position of each pixel position in the transformed dial image in the dial image according to the inverse transformation parameters; determining the pixel value of each pixel position in the transformed dial image according to the pixel value at each mapping position; and determining the transformed dial image according to the pixel value of each pixel position in the transformed dial image.
[0083] First, the server performs a matrix inversion operation on the perspective transformation parameters (i.e., the homography matrix) to obtain the corresponding inverse matrix, which is the inverse transformation parameter. The inverse transformation parameter describes the coordinate transformation relationship from the target template image plane (i.e., the corrected standard frontal view plane) back to the original dial image plane.
[0084] Then, the server determines the size of the transformed dial image, ensuring it matches the target template image size to guarantee spatial resolution alignment between the corrected dial and the standard template. The server iterates through each pixel in the transformed dial image. For each pixel, it multiplies its homogeneous coordinates with the inverse matrix represented by the inverse transformation parameters, obtaining a mapping result in homogeneous coordinate form through matrix multiplication. The server then converts the homogeneous coordinates to Cartesian coordinates to obtain the corresponding mapped position of that pixel in the original dial image. This can be represented by the following formula:
[0085]
[0086] In the formula, Mapping function for homogeneous coordinates to Cartesian coordinates; Homogeneous coordinate vector It is a non-zero scaling factor; The three components of homogeneous coordinates; , Transformed Cartesian coordinates and Quantity.
[0087] From the above formula, we can see that homogeneous coordinates Points on a two-dimensional plane are scaled by adding a scale factor. Upgrading to a three-dimensional homogeneous space unifies translation and perspective projection transformations, which were previously unrepresentable by linear matrices, into linear matrix operations. Transformation functions This completes the mapping from homogeneous space back to two-dimensional Cartesian space, where the scale factor... The existence of this property ensures that the transformed coordinates are scale invariant.
[0088] Next, the server determines the pixel value of each pixel position in the transformed dial image based on the pixel values at each mapped position. For a target pixel position in the transformed dial image, its corresponding target mapped position usually does not coincide with the integer pixel grid of the original dial image. Therefore, the server uses a bilinear interpolation method to sample pixel values from multiple integer pixels adjacent to the target mapped position in the original dial image. Specifically, the server calculates a weighting coefficient for multiple adjacent pixels based on the decimal part of the coordinates of the target mapped position. The larger the decimal part of the coordinates, the greater the contribution of adjacent pixels in that direction to the target pixel value. The server sums the pixel values of multiple adjacent pixels according to the above weighting coefficients to obtain the interpolated pixel value at the target mapped position, and assigns the interpolated pixel value to the corresponding target pixel position in the transformed dial image. The server iterates through all pixel positions in the transformed dial image, performing the above mapping and interpolation operations for each position, thereby completing the determination of all pixel values.
[0089] Finally, the server constructs and outputs the transformed dial image based on the pixel values corresponding to each pixel position in the transformed dial image. The transformed dial image has the same size as the target template image, and the dial plane is parallel to the imaging plane, presenting a standard frontal view.
[0090] In this embodiment, the reverse mapping strategy (i.e., traversing every pixel in the output image and calculating its mapping position in the original image in reverse) can effectively avoid the problems of empty pixels and pixel overlap in the output image compared with the forward mapping strategy, thus ensuring the integrity and continuity of the transformed dial image.
[0091] In an exemplary embodiment, determining the pointer position in the transformed dial image and determining the dial reading of the dial to be read based on the pointer position and the scale reference parameters corresponding to the target template image includes: identifying the pixel region where the dial pointer is located from the transformed dial image and determining the pixel position belonging to the dial pointer to obtain a pointer segmentation result; determining the pointer tip position and pointer skeleton direction of the dial pointer in the transformed dial image based on the pointer segmentation result; and determining the dial reading corresponding to the dial image based on the pointer tip position, pointer skeleton direction, and the scale reference parameters corresponding to the target template image.
[0092] The pointer segmentation result refers to the binary mask data that identifies which pixels in the transformed dial image belong to the pointer region. Pixels belonging to the pointer are marked as foreground, and pixels not belonging to the pointer are marked as background. The scale reference parameters include the dial scale start pixel position, scale end pixel position, lower limit of the range, and upper limit of the range.
[0093] First, the server identifies the pixel region containing the dial pointer from the transformed dial image and determines the pixel positions belonging to the dial pointer, thus obtaining the pointer segmentation result. Specifically, the server inputs the transformed dial image into a pre-trained instance segmentation model for forward inference computation. The instance segmentation model employs an encoder-decoder architecture. The encoder performs multi-level downsampling on the input image to extract deep semantic features, and the decoder progressively upsampling the deep features to restore the original image resolution, outputting a pointer region probability prediction map with the same size as the input image. The value of each pixel position in the probability prediction map represents the probability that the corresponding position belongs to the pointer region, ranging from 0 to 1. The server segments the probability prediction map using a preset binarization threshold. Pixel positions with a probability value greater than or equal to the threshold are identified as pointer regions and set as foreground values, while pixel positions with a probability value less than the threshold are identified as non-pointer regions and set as background values, thereby obtaining the pointer segmentation result.
[0094] Then, the server extracts the pointer's geometric features based on the binary mask in the pointer segmentation result. First, the server performs skeleton thinning on the binary mask, iteratively stripping away edge pixels to shrink the pointer region into a single-pixel-wide skeleton line. This skeleton line preserves the pointer region's topology and direction of extension, containing multiple consecutive skeleton points. The server identifies the end of the skeleton line furthest from the dial's center as the pointer tip position. Specifically, it calculates the Euclidean distance between each skeleton point and the center position in the transformed dial image, as indicated by the scale reference parameters, and identifies the skeleton point with the largest distance as the pointer tip position. Based on the coordinates of several consecutive skeleton points near the pointer tip position on the skeleton line, the server calculates the direction angle of the local line segment formed by these skeleton points, using this direction angle as the pointer skeleton direction, representing the direction the pointer is pointing.
[0095] Finally, the server determines the dial reading based on the pointer tip position, pointer skeleton direction, and the scale reference parameters corresponding to the target template image. The server maps the starting and ending pixel positions of the scale in the target template image to the coordinate space of the transformed dial image, obtaining the corresponding starting and ending angles on the transformed dial image. Simultaneously, the center position of the dial circle marked in the target template image is used as the reference origin for the angles. The server connects the pointer tip position to the center position and calculates the deflection angle of the line relative to the horizontal direction, which is the actual pointing angle of the pointer on the dial, i.e., the pointer pointing angle. Based on the proportional relationship between the pointer pointing angle, the starting angle, and the ending angle, and combined with the lower and upper limits of the range, the server calculates the current reading using linear interpolation. Specifically, the ratio of the difference between the pointer pointing angle and the starting angle to the total difference between the starting and ending angles, multiplied by the difference between the upper and lower limits of the range, and then added to the lower limit of the range, yields the current pointer reading. For dials with non-linear scales, multiple sets of key intermediate scale points and their corresponding range values are pre-marked in the target template image. The server calculates the current reading using piecewise linear interpolation based on the position of the pointer's pointing angle in each segment interval, thereby adapting to meters with different scale distribution types.
[0096] In this embodiment, pointer segmentation and geometric analysis are performed on the standard viewpoint dial image after perspective correction, avoiding the problems of inaccurate center positioning and amplified pointer angle deviation in distorted images, and achieving high-precision pointer position determination and reading calculation.
[0097] In an exemplary embodiment, determining the dial reading corresponding to the dial image based on the pointer tip position, pointer skeleton direction, and scale reference parameters corresponding to the target template image includes: determining the pointer deflection angle based on the pointer tip position, pointer skeleton direction, and the center position of the transformed dial image; determining the scale start angle and scale end angle based on the scale start position and scale end position in the scale reference parameters corresponding to the target template image, and the center position of the transformed dial image; determining the range values corresponding to the scale start angle and scale end angle based on the scale reference parameters corresponding to the target template image; and determining the dial reading corresponding to the dial image based on the pointer deflection angle, scale start angle, scale end angle, and the range values corresponding to the scale start angle and scale end angle.
[0098] The center position of the transformed dial image is derived from the dial center coordinates marked in the scale reference parameters of the target template image. These center coordinates are mapped to the coordinate space of the transformed dial image during the inverse perspective transformation process, along with the inverse transformation parameters. The pointer deflection angle refers to the angle formed by the direction vector from the dial center to the pointer tip relative to a preset zero-degree reference direction (e.g., horizontal to the right).
[0099] First, the server constructs a pointer direction vector, starting from the center of the transformed dial image and ending at the pointer tip. Then, it calculates the angle between this vector and a preset zero-degree reference direction to obtain the pointer deflection angle. Simultaneously, the server verifies the calculation result using the pointer skeleton direction. If the deviation between the direction determined by the pointer tip position and the pointer skeleton direction is within a preset range, the pointer deflection angle is confirmed as valid. If the deviation exceeds the preset range, the server corrects the pointer deflection angle using a weighted average of the pointer tip position and the skeleton direction to eliminate tip positioning errors caused by incomplete pointer segmentation.
[0100] Then, the server determines the scale start angle and scale end angle based on the scale start and end positions in the scale reference parameters corresponding to the target template image, and the center position of the transformed dial image. The scale reference parameters corresponding to the target template image indicate the scale start and end pixel positions, which are also mapped to the coordinate space of the transformed dial image during the inverse perspective transformation. The server constructs two direction vectors, using the scale start and end positions as endpoints and the center position of the transformed dial image as the starting point, and calculates the angles between these two vectors relative to the preset zero-degree reference direction to obtain the scale start angle and scale end angle. The scale start angle and scale end angle correspond to the spatial orientation of the lower and upper limit scale lines of the dial range on the dial circumference, respectively.
[0101] Next, the server determines the range values corresponding to the scale start angle and scale end angle based on the scale reference parameters corresponding to the target template image. The scale reference parameters corresponding to the target template image include a lower range limit and an upper range limit. The lower range limit is the physical reading corresponding to the scale start position, and the upper range limit is the physical reading corresponding to the scale end position. The server uses the lower range limit as the range value corresponding to the scale start angle and the upper range limit as the range value corresponding to the scale end angle, thereby establishing a mapping relationship between the dial's spatial angle and the physical reading.
[0102] Finally, the server determines the direction of the scale area, i.e., whether the pointer moves clockwise or counterclockwise from the starting to the ending scale on the dial, and determines the method for calculating the angle difference accordingly. When the scale area extends clockwise, the server calculates the clockwise offset of the pointer deflection angle relative to the starting angle of the scale, as well as the total clockwise angle span between the starting and ending angles. Then, it calculates the proportion of the clockwise offset to the total angle span, multiplies this proportion by the difference between the upper and lower limits of the dial's range, and adds the lower limit to obtain the dial reading. When the scale area extends counterclockwise, the server uses the same proportional interpolation calculation with the counterclockwise offset and angle span. For dials with non-linear scales, the scale reference parameters corresponding to the target template image also include one or more intermediate scale key points and their corresponding angle and range values. The server first identifies which segment interval the pointer deflection angle falls within between two adjacent scale key points, and then uses a linear interpolation method to calculate the dial reading within that segment interval.
[0103] In this embodiment, the nonlinear scale distribution of the circular dial is transformed into a linear interpolation problem in angle space, which is compatible with both linear and nonlinear dial types. This avoids the systematic deviation caused by directly reading the pointer position in distorted images and ensures the accuracy and universality of the reading recognition results.
[0104] In an exemplary embodiment, determining the pointer deflection angle based on the pointer tip position, the pointer skeleton direction, and the center position of the transformed dial image includes: determining the orientation of the pointer tip relative to the center of the circle based on the center position of the transformed dial image and the pointer tip position, obtaining a first azimuth angle; determining the direction of the pointer skeleton relative to the center of the circle based on the pointer skeleton direction, obtaining a second azimuth angle; performing a consistency check on the first azimuth angle and the second azimuth angle to obtain a consistency check result between the first azimuth angle and the second azimuth angle; and determining the pointer deflection angle based on the consistency check result between the first azimuth angle and the second azimuth angle.
[0105] First, the server constructs a polar coordinate system using the center position of the transformed dial image as the origin and a preset zero-degree reference direction (e.g., horizontal to the right) as the angular reference axis. The server calculates the vector from the center position to the pointer tip position. In the polar coordinate system, the angle between this vector and the angular reference axis is the azimuth angle of the pointer tip relative to the center position, and the server records this azimuth angle as the first azimuth angle.
[0106] Then, the server determines the azimuth angle pointed to by the pointer skeleton line extending outward from the center of the circle based on the angle between the pointer skeleton direction and the preset zero-degree reference direction, and records this azimuth angle as the second azimuth angle.
[0107] Next, the server calculates the angle difference between the first and second azimuth angles, takes the absolute value of the angle difference, and compares it with a preset consistency angle threshold. When the angle difference is less than or equal to the preset consistency angle threshold, the consistency check is passed, indicating that the first azimuth angle determined by the pointer tip position and the second azimuth angle determined by the pointer skeleton direction match, and the pointer tip position and skeleton direction both point to the same direction, indicating that the pointer geometric feature extraction result is reliable. When the angle difference is greater than the preset consistency angle threshold, the consistency check is failed, indicating that there is an inconsistency between the pointer tip position and the skeleton direction, and there may be errors in the pointer segmentation or skeleton extraction process.
[0108] Finally, when the consistency check passes, the server fuses the first and second azimuth angles, specifically calculating their weighted average as the final pointer deflection angle. The weighting coefficients are pre-set based on the confidence levels of the first and second azimuth angles. When the consistency check fails, the server further determines whether the angle difference exceeds a preset failure threshold. If the angle difference does not exceed the failure threshold but exceeds the consistency angle threshold, the server uses the first azimuth angle as the pointer deflection angle, as the pointer tip position directly determines the pointer's pointing end, making its positioning accuracy more critical. If the angle difference exceeds the preset failure threshold, it indicates a serious error in the pointer segmentation result. The server abandons the current pointer deflection angle calculation, triggers a segmentation anomaly alarm, and terminates the current dial reading recognition process.
[0109] In this embodiment, by utilizing two types of geometric information—the pointer tip position and the pointer skeleton direction—a consistency verification and fusion decision mechanism is used to suppress the impact of incomplete pointer segmentation or skeleton extraction noise on the accuracy of deflection angle calculation.
[0110] In an exemplary embodiment, determining the pointer position in the transformed dial image includes: mapping the pixels in the scale area of the transformed dial image to a rectangular coordinate space according to their respective polar coordinate parameters, with the center of the circle of the transformed dial image as the origin, to obtain a polar coordinate unfolded dial image; in the polar coordinate unfolded dial image, determining the number of pixels belonging to the pointer at each position perpendicular to the unfolding direction along the unfolding direction, and obtaining a statistical result; and determining the position of the pointer in the unfolding direction based on the position with the most pixels in the statistical result.
[0111] First, the server reads the dial's center position and the inner and outer diameters of the effective scale area from the scale reference parameters corresponding to the target template image. Using the center as the origin of polar coordinates and a preset zero-degree direction as the starting angle direction, the server divides the dial's annular scale area into multiple grid units arranged radially and angularly. For each grid unit, the server samples the corresponding pixel value in the original transformed dial image based on its radial distance and angular position, and fills this pixel value into the corresponding row and column positions in the rectangular coordinate space. Here, the row direction of the rectangular coordinate space corresponds to the radial distance in polar coordinates, and the column direction corresponds to the angular direction in polar coordinates, thus unfolding the annular scale area into a rectangular two-dimensional image. After this polar coordinate unfolding process, the original circumferentially distributed scale lines and pointers on the dial are transformed into a horizontally linearly arranged pattern in the rectangular coordinate space, with the pointer appearing as a horizontally extending strip-shaped area in the unfolded image.
[0112] Then, the server defines the column direction (i.e., the angular direction) and the row direction (i.e., the radial direction) as the unfolding direction in the polar coordinate unfolded dial image, and defines the vertical direction as the vertical direction. For each column position in the unfolding direction, the server counts the number of pixels belonging to the pointer region among all pixels in the vertical direction at that column position. The determination of pixels belonging to the pointer region is based on the pointer segmentation result, that is, the server maps the pointer binary mask obtained in the previous step to the polar coordinate unfolded image space according to the same polar coordinate transformation parameters, and then counts the number of pixels with mask values in the foreground for each column position in the vertical direction. The server records the statistical values corresponding to each column position, forming a pixel number distribution curve along the unfolding direction.
[0113] Next, the server iterates through all statistical values on the pixel distribution curve, finding the peak position with the largest value. The column coordinates corresponding to the peak position are used as the pointer's position in the unfolding direction. Since the pointer appears as a strip-shaped region extending along the unfolding direction in the polar coordinate unfolded image, the span of the pointer in this direction corresponds to the angular range occupied by the pointer in the original circular dial. Therefore, the position of the pointer tip corresponds to the termination boundary position of the pointer region in the unfolded image in the unfolding direction. Based on the pointer's geometric characteristics, the server further extracts the pointer tip position from the pointer's position in the unfolding direction. Specifically, in the unfolding direction, from the leading or trailing boundary of the pointer region, the boundary position furthest from the starting mark is determined according to the pointer's pointing direction as the pointer tip's position in the unfolding direction.
[0114] In this embodiment, the non-linear scale area of the circular dial is transformed into a linear rectangular space through polar coordinate transformation. This transforms the determination of the pointer position from angle calculation on a two-dimensional plane to peak detection in a one-dimensional unfolding direction, avoiding the influence of the center positioning error when directly calculating the angle on the distortion-corrected image.
[0115] In an exemplary embodiment, before extracting the feature vector of the dial image, the method further includes: obtaining a standard dial template image; cropping the standard dial template image to obtain a template dial area image; determining the scale reference parameters in the template dial area image; the scale reference parameters include the scale start position, the scale end position, and the range values corresponding to the scale start position and the scale end position, respectively; extracting the feature vector of the template dial area image as the template feature vector corresponding to the standard dial template image; and constructing a template feature library based on the standard dial template image, the template feature vector, and the scale reference parameters.
[0116] First, maintenance personnel use industrial cameras or high-definition imaging equipment to capture standard, distortion-free, high-resolution images of various analog meters from the normal direction facing the dial, ensuring that the dial markings, numerals, and pointer outlines are clear and complete. The server receives and stores these captured raw standard dial images as input data for subsequent processing.
[0117] Then, the server uses image annotation tools or automatic segmentation algorithms to accurately crop the effective area of the pure watch face from the original standard watch face image, removing the outer border, casing, and background interference areas. Simultaneously, the server uniformly scales the cropped template watch face area image to a preset size to adapt to the input requirements of the subsequent feature extraction network, ensuring that the template watch face area image has a uniform spatial resolution, thereby maintaining consistent expression of feature vectors.
[0118] Next, the server uses an image annotation tool to mark the start and end pixels of the scale in the template dial area image, corresponding to the coordinate positions of the lower and upper limit scale lines of the dial range in the image, respectively. The server also marks the dial's center reference point and the inner and outer diameters of the effective scale area, used to determine the expansion range during subsequent polar coordinate expansion. Based on the marked start and end positions of the scale, and the actual lower and upper limit values of the meter's range entered by the maintenance personnel, the server establishes a correspondence between the start and end positions of the scale and between the end and upper limits of the scale. These coordinate and numerical parameters are then stored in a structured data format to form the scale reference parameters for the template dial area image.
[0119] Then, the server calls the pre-trained feature extraction model, inputs the cropped template dial area image into the feature extraction model for forward inference calculation. The feature extraction model extracts high-dimensional global feature representations from the template dial area image through multi-layer convolution and pooling operations, and finally outputs a fixed-dimensional feature vector to represent the appearance, texture and structural features of the template dial area image.
[0120] Finally, the server correlates the template feature vector of the standard dial template image, the image data of the standard dial template image, and the scale reference parameters corresponding to the standard dial template image to obtain the template registration record. The server performs the above operations of acquisition, cropping, parameter calibration, feature extraction, and associated storage for all meter types that need to be identified in the substation, and persistently stores the template registration records of all meter types in the server's database to obtain the template feature library.
[0121] In this embodiment, adding a new meter type only requires supplementing the corresponding template registration record without retraining the recognition model, which reduces the engineering cost and deployment cycle of the system in the process of adapting to multiple types of meters.
[0122] Figure 4 For a complete flowchart of a dial reading recognition method in another embodiment, please refer to [link / reference]. Figure 4In a detailed embodiment, firstly, during the offline template registration phase, maintenance personnel collect standard, distortion-free, high-definition images of various substation pointer meters. Using image annotation tools, they crop out the effective area of the pure dial to obtain standard template images. Each template image is then annotated with the scale start and end pixels, the corresponding lower and upper range values, the dial center reference point, and the inner and outer radii of the effective scale area. All coordinates and range parameters are associated and stored as a template-specific annotation file. Next, the Fast-ReID feature extraction model (with a ResNet-50-ibn backbone network, jointly optimized using Circle Loss and Center Loss), fine-tuned with a substation-specific dataset, is called to extract a 2048-dimensional feature vector from each template image. This feature vector is then associated with the corresponding template image and scale reference parameters to construct and persistently store the substation meter template feature library (i.e.,...). Figure 4 Template libraries and template vector libraries in [the library / database].
[0123] Then, in the online real-time reading recognition stage, the original images of the substation inspection captured by the inspection robot or fixed camera (i.e., Figure 4 The original image (in the image above) is used to perform dial target detection on the inspection original image. Unreliable detection boxes are filtered out based on a pre-set confidence threshold, and the dial image to be read is cropped from the inspection original image. Then, the Fast-ReID model with the same weights as the template registration stage is called to extract the feature vectors of the dial image. The cosine similarity algorithm is then used to calculate the similarity score between each feature vector and the feature vectors of each template in the template feature library. The template with the highest score is selected as the target template image (i.e., the target template image). Figure 4 (The template corresponding to the target template image is used); if the highest similarity is lower than the preset threshold, it is determined to be an unknown meter and the process is terminated; otherwise, the scale reference parameters associated with the target template image are read.
[0124] Next, the dial image to be read and the target template image are input into a LoFTR feature extraction network based on the Transformer architecture. The feature extraction network extracts multi-scale feature maps of the two images through a CNN backbone with shared weights. After adding two-dimensional sinusoidal positional encoding to the multi-scale feature maps, they are sequentially passed through a multi-layer self-attention module of the feature extraction network to achieve global context awareness of a single image. After the multi-scale feature maps are processed by the multi-layer self-attention module of the feature extraction network, a bidirectional cross-attention module is used to establish the pixel-level cross-domain feature correspondence between the two images. Then, a matching score matrix is constructed based on optimal transfer theory, and the double random matching matrix is solved iteratively using the Sinkhorn algorithm to select multiple sets of high-precision matching point pairs with confidence scores higher than a threshold. On this basis, the RANSAC algorithm combined with the DLT method is used to iteratively select interior points from multiple sets of matching point pairs, eliminate mismatched exterior points, and solve for the optimal perspective transformation parameters (i.e., homography matrix) of the dial image relative to the target template image. Based on the perspective transformation parameters, the inverse transformation parameters are obtained. A reverse mapping strategy is used to traverse each pixel position in the transformed dial image. The mapping position in the original dial image is calculated using the inverse matrix. Pixel values are then sampled using bilinear interpolation to finally reconstruct a distortion-free, standard front-view transformed dial image (e.g., ...). Figure 4 The result after the middle dial is corrected is shown in the image.
[0125] After obtaining the transformed dial image, the YOLOv11-seg instance segmentation model is used to perform pixel-level segmentation of the pointer region, outputting a binary pointer mask. Morphological opening and closing operations are then employed for joint optimization to eliminate discrete noise and voids. Based on the optimized mask, the pointer tip position and pointer skeleton direction are extracted. Simultaneously, the pointer deflection angle is calculated by combining the center position of the circle obtained from the scale reference parameters in the transformed dial image. To eliminate the influence of the non-linear distribution of the circular dial scale, the effective scale region of the transformed dial image is expanded into a rectangular coordinate space with the center as the origin of polar coordinates, based on radial distance and effective angle intervals. This yields the polar coordinate expanded dial image. The starting and ending positions of the scale marked on the template are then synchronously mapped to polar coordinate angles, completing the linear alignment of the scale reference. In the unfolded rectangular image, the pointer mask is integrally projected along the radial vertical direction to obtain the pointer horizontal position distribution curve. The horizontal coordinate corresponding to the projection peak is taken as the pointer position in the unfolding direction. Finally, combined with the pixel positions of the template scale start and end points and the corresponding lower and upper range values, the final dial reading of the dial to be read is calculated by linear interpolation (piecewise linear interpolation is used for non-linear scale dials).
[0126] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0127] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores standard dial template images, template feature vectors, and corresponding scale reference parameters. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a dial reading recognition method.
[0128] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0130] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0131] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0132] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for recognizing dial readings, characterized in that, include: Acquire the dial image of the watch face to be read, and extract the feature vector of the dial image; The feature vector is matched with each template feature vector in the preset template feature library to find the target template image that matches the dial image; the template feature library stores at least one standard dial template image, a template feature vector corresponding to each standard dial template image, and a scale reference parameter corresponding to each standard dial template image; The dial image and the target template image are input into a feature extraction network to obtain multi-scale feature maps of the dial image and the target template image. For each location in the multi-scale feature maps of the dial image and the target template image, feature information from other locations is fused to obtain a first global feature of the dial image and a second global feature of the target template image. The multi-scale feature maps include at least a low-resolution feature map and a high-resolution feature map. The low-resolution feature map is used to represent the global structural information and global semantic information of the corresponding image, while the high-resolution feature map is used to represent the local structural information and local semantic information of the corresponding image. Cross-attention matching is performed on the first global feature and the second global feature to obtain multiple sets of matching point pairs; Based on the multiple sets of matching point pairs, the perspective transformation parameters of the dial image relative to the target template image are determined; According to the perspective transformation parameters, the dial image is subjected to inverse perspective transformation to obtain the transformed dial image; The pointer position is determined in the transformed dial image, and the dial reading of the dial to be read is determined according to the pointer position and the scale reference parameters corresponding to the target template image.
2. The method according to claim 1, characterized in that, Cross-attention matching is performed on the first global feature and the second global feature to obtain multiple sets of matching point pairs, including: Cross-attention processing is performed on the first global feature and the second global feature to establish the feature correspondence between the dial image and the target template image; Based on the feature correspondence, matching feature points are selected from the dial image and the target template image to obtain multiple sets of matching point pairs.
3. The method according to claim 2, characterized in that, Based on the feature correspondence, matching feature points are selected from the dial image and the target template image to obtain multiple sets of matching point pairs, including: Based on the feature correspondence, a matching score matrix is constructed; the element values in the matching score matrix are used to characterize the degree of matching between the feature points in the dial image and the feature points in the target template image. Based on the values of each element in the matching score matrix, the matching score matrix is optimized to obtain an optimized matching score matrix; the optimized matching score matrix is used to establish a globally optimal one-to-one correspondence between the feature points of the dial image and the feature points of the target template image; Based on the optimized matching score matrix, mutually matching feature points are selected from the dial image and the target template image to obtain the multiple sets of matching point pairs.
4. The method according to claim 1, characterized in that, Based on the multiple sets of matching point pairs, the perspective transformation parameters of the dial image relative to the target template image are determined, including: Select one set of matching point pairs from the multiple sets of matching point pairs; Based on the selected matching point pairs, determine the initial perspective transformation parameters; Based on the initial perspective transformation parameters, each matching point in the dial image is projected onto the coordinate space of the target template image to obtain the projection position corresponding to each matching point. Calculate the positional error between the projection position of each matching point and the corresponding matching point in the target template image; The matching point pairs whose position error is less than a preset error threshold are determined as valid matching point pairs; The initial perspective transformation parameters are updated based on the number of valid matching point pairs; Return to the step of selecting a set of matching point pairs from the multiple sets of matching point pairs, until the preset iteration stop condition is met, and obtain the perspective transformation parameters.
5. The method according to claim 1, characterized in that, Based on the perspective transformation parameters, the dial image is subjected to inverse perspective transformation to obtain the transformed dial image, including: Obtain the inverse transformation parameters of the perspective transformation parameters; Based on the inverse transformation parameters, determine the mapping position of each pixel position in the transformed dial image within the dial image; Based on the pixel values at each of the mapped positions, determine the pixel values at each pixel position in the transformed dial image; The transformed dial image is determined based on the pixel values of each pixel position in the transformed dial image.
6. The method according to claim 1, characterized in that, Determine the pointer position in the transformed dial image, and determine the dial reading of the dial to be read based on the pointer position and the scale reference parameters corresponding to the target template image, including: The pixel region where the dial pointer is located is identified from the transformed dial image, and the pixel position belonging to the dial pointer is determined to obtain the pointer segmentation result; Based on the pointer segmentation results, determine the pointer tip position and pointer skeleton direction of the dial pointer in the transformed dial image; The dial reading corresponding to the dial image is determined based on the position of the pointer tip, the direction of the pointer skeleton, and the scale reference parameters corresponding to the target template image.
7. The method according to claim 6, characterized in that, Based on the position of the pointer tip, the direction of the pointer skeleton, and the scale reference parameters corresponding to the target template image, the dial reading corresponding to the dial image is determined, including: The pointer deflection angle is determined based on the position of the pointer tip, the direction of the pointer skeleton, and the center position of the transformed dial image. Based on the scale start position and scale end position in the scale reference parameters corresponding to the target template image, and the center position of the transformed dial image, the scale start angle and scale end angle are determined respectively. Based on the scale reference parameters corresponding to the target template image, determine the range values corresponding to the scale start angle and the scale end angle respectively; The dial reading corresponding to the dial image is determined based on the pointer deflection angle, the scale start angle, the scale end angle, and the range values corresponding to the scale start angle and the scale end angle, respectively.
8. The method according to claim 7, characterized in that, The pointer deflection angle is determined based on the position of the pointer tip, the direction of the pointer skeleton, and the center position of the transformed dial image, including: Based on the center position of the transformed dial image and the position of the pointer tip, the orientation of the pointer tip relative to the center of the circle is determined, and the first azimuth angle is obtained. Based on the direction of the pointer skeleton, the pointing of the pointer skeleton relative to the center of the circle is determined, and the second azimuth angle is obtained; A consistency check is performed on the first azimuth angle and the second azimuth angle to obtain the consistency check result between the first azimuth angle and the second azimuth angle; The pointer deflection angle is determined based on the consistency verification result between the first azimuth angle and the second azimuth angle.
9. The method according to claim 1, characterized in that, Determining the pointer position in the transformed dial image includes: Using the center of the transformed dial image as the origin, the pixels in the scale area of the transformed dial image are mapped to a rectangular coordinate space according to their respective polar coordinate parameters to obtain the dial image after polar coordinate expansion. In the polar coordinate unfolded dial image, the number of pixels belonging to the pointer at each position perpendicular to the unfolding direction is determined sequentially along the unfolding direction to obtain statistical results; The position of the pointer in the unfolding direction is determined based on the position with the most pixels in the statistical results.
10. The method according to any one of claims 1-9, characterized in that, Before extracting the feature vector of the dial image, the following steps are also included: Obtain a standard watch face template image; The standard dial template image is cropped to obtain a template dial area image; Determine the scale reference parameters in the template dial area image; the scale reference parameters include the scale start position, the scale end position, and the range values corresponding to the scale start position and the scale end position respectively; Extract the feature vector of the template dial area image and use it as the template feature vector corresponding to the standard dial template image; The template feature library is constructed based on the standard dial template image, the template feature vector, and the scale reference parameters.