Industrial pointer type disc instrument intelligent reading method for complex environment
Through deep learning models based on YOLOv10 and IPT, STN, PaddleOCR, and Deeplabv3+, the problem of instrument reading in complex environments is solved, and the fast and high-precision intelligent reading of industrial pointer disc instruments is achieved, which improves reading speed and accuracy and reduces costs.
Patent Information
- Application Number
- CN202510363222.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Traditional industrial meter reading methods are low efficiency and accuracy, and the instrument reading is difficult in complex environments. The existing deep learning methods have problems such as weak feature representation ability, many model parameters, and long inference time.
The deep learning model based on YOLOv10 is used for dashboard detection, combined with the IPT model for image repair and STN algorithm for image correction, used PaddleOCR to identify the range and unit of measurement, combined with Deeplabv3+ for scale segmentation, detected the pointer position through Hough transformation, and finally used the center angle algorithm for reading.
It realizes fast and high-precision instrument readings in complex environments, improves reading speed and accuracy, and reduces human resource consumption and implementation costs.
Smart Images

Figure CN120299018A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically provides an intelligent reading method for industrial pointer-type dial instruments facing complex environments. Background Art
[0002] Traditional industrial meter reading requires personnel to go to the site for meter reading, which consumes manpower and material resources, is prone to data entry errors, and poses safety risks. Therefore, traditional industrial meter reading solutions cannot meet the requirements of modern industries for accurate metering and safety management. Existing industrial pointer-type dial instrument reading methods include: 1. Traditional manual observation. This method is greatly affected by subjective factors and objective conditions and has low efficiency. 2. Using traditional image processing technology. These methods have low accuracy, usually require manual adjustment of multiple parameters, and have poor versatility. 3. Based on deep learning technology. These methods have brought a double improvement in reading accuracy and speed.
[0003] However, there are still many challenges in instrument reading in complex environments, mainly including: 1. Industrial instrument dials are small, pointers are thin, and it is difficult to detect instruments. 2. The dial is affected by complex environmental interferences such as oil stains, cracks, and tilts, making it difficult to read the instrument. 3. Existing instrument reading methods based on deep learning have problems such as weak feature representation ability, many model parameters, long inference time, etc., and the instrument reading accuracy is low. Summary of the Invention
[0004] In view of the above problems of the prior art, the present invention provides an intelligent reading method for industrial pointer-type dial instruments facing complex environments, which can achieve fast and high-precision recognition of instrument readings.
[0005] To achieve the above object, the present invention proposes an intelligent reading method for industrial pointer-type dial instruments facing complex environments, including the following steps:
[0006] S1. Detect the instrument panel included in the input image;
[0007] S2. Denoise the instrument panel image;
[0008] S3. Identify the instrument range and measurement unit based on the instrument panel image;
[0009] S4. Identify the center of the instrument panel, zero scale, maximum scale, and the scale pointed by the pointer based on the instrument panel image, and construct a simplified instrument panel schematic diagram;
[0010] S5. Read the instrument according to the simplified instrument panel schematic diagram, instrument range, and measurement unit.
[0011] Preferably, in S1, the detection of the dashboard included in the input image is divided into five stages: acquisition, sample annotation, data input, model optimization, and model inference. The specific steps are as follows:
[0012] S11. Acquisition: Deploy a camera near the industrial pointer-type dial instrument to collect instrument pictures;
[0013] S12. Sample annotation: Manually annotate the instrument pictures using the labellmg software;
[0014] S13. Data input: Input the data set into the deep learning network model based on YOLOv10 to train the model;
[0015] S14. Model optimization: Optimize through the loss function to obtain the optimal weights;
[0016] S15. Model inference: Input the instrument image into the trained model. The model outputs the detection box parameters and categories of the instrument, calls the function to crop out the detection box, intercept the dashboard included in the input image, and store it in the corresponding folder.
[0017] Preferably, in S14, the loss function is Shape-IoU, which can be expressed as:
[0018] L S hape-Io U = 1 - IoU + distance shape + 0.5′Ω shape ;
[0019] In the formula, IoU is the intersection over union of the GT box and the anchor box, and distance shape is the distance between the GT box and the anchor box.
[0020] Preferably, in S2, the specific steps for denoising the dashboard image are to take different operations according to the instrument categories output in the inference stage of the YOLOv10 model: If it is greasy or broken, first perform image restoration and then image correction; if it is normal, directly perform image correction.
[0021] Preferably, in S2, the IPT model is used for image restoration. The model consists of four parts: a head that extracts features from the dashboard pictures, an encoder and a decoder that recover the lost information, and a tail that maps the features to the restored image. The objective function of IPT is defined as:
[0022] L IPT = λ·L contrastive + L supervised ;
[0023] Wherein, L contrastive is the contrast loss, and L supervised is the supervision loss.
[0024] Preferably, in S2, the image correction is based on the STN algorithm to perform image transformation on the dashboard image; the Spatial Transformer Network (STN) converts the dashboard picture into a front-facing dashboard picture, and it includes three steps: parameter prediction Localisation net, coordinate mapping Grid generator, and pixel sampling Sampler. The specific operations are as follows:
[0025] S21. Localisation net calculates the parameters required for spatial transformation: Input a Feature map: U ∈ R H×W×C , and after several convolutional and fully connected operations, a regression layer is used to obtain the transformation parameter θ. The formula is as follows:
[0026] θ = f loc (U);
[0027] Wherein, θ is the transformation parameter, the dimension of θ depends on the specific transformation type selected by the network, U ∈ R H×W×C is the input feature map, and floc(·) is the localization network, and the localization network floc(·) can take any form;
[0028] S22. Grid generator uses the calculated θ to perform corresponding spatial transformation on the Feature map. The formula is as follows:
[0029]
[0030] In the formula, represents the target pixel position on the output feature map, T θ (G i ) represents the applied transformation, G i represents the regular network, A θ represents the transformation matrix, represents the sampling position on the corresponding input feature map;
[0031] S23. Sampler extracts pixel values from the input feature map (U) according to the positions in the transformed sampling grid T θ (G i ) to generate the output feature map (V); each obtained position in the sampling grid indicates a specific sampling point on the input feature map (U).
[0032] Preferably, in S3, the specific steps for identifying the instrument range and measurement unit based on the dashboard image are as follows:
[0033] S31. Call the trained text recognition model PaddleOCR, configure the path frame of the dashboard image to be recognized, store the recognition result in result, and extract the pure text part stored in result and store it in text;
[0034] S32. Match text according to the characteristic information of the instrument range and measurement unit, and classify and categorize the two. The formula is as follows:
[0035]
[0036] In the formula, Ks represents the candidate set of instrument ranges, and Us represents the candidate set of measurement units.
[0037] S33. Select the one with the largest value in the Ks set as the instrument range K, and select the one that matches the actual situation in the Us set as the measurement unit U.
[0038] Preferably, in S4, based on the dashboard image, recognize the center of the dashboard, the zero scale, the maximum scale, and the scale pointed by the pointer, and the specific steps for constructing a simplified dashboard schematic diagram are as follows:
[0039] S41. Use the trained semantic segmentation model Deeplabv3+ to segment the scales in the dashboard image, and read the segmentation mask. The specific operation steps are as follows:
[0040] After the output layer of DeepLabv3+, cross-entropy loss is used to calculate the loss between the predicted class map and the true label, and according to the calculated loss, the weights of the network are updated through the backpropagation algorithm. The cross-entropy loss function formula is:
[0041] L = -∑ i y i log(p i ) - (1 - y i )log(1 - p i );
[0042] In the formula, p is the probability that the model predicts to belong to the instrument scale category, and y is the true label. Among them, y = 1 means that the pixel belongs to the instrument scale, and y = 0 means that the pixel does not belong to the instrument scale;
[0043] S42. Fit the circle where the scale is located, and calculate the center coordinates O(x0, y0). Use Thomas corner detection to obtain the two endpoints A(x1, y1) and B(x2, y2) of the fitting circular curve of the scale. Connect the two endpoints and the center respectively to obtain the zero scale line and the maximum scale line; among them, the center of the dashboard is the center coordinates, and perform ellipse fitting and Thomas corner detection on the picture grid_image and output the processed picture;
[0044] S43. Determine the position of the pointer line based on the Hough transform line detection. Among them, select the longest line object as the pointer and store the position information of the pointer. The specific steps to determine the position of the pointer line are as follows:
[0045] Associate each line in the image with the Hough space of a pair of parameters (r,θ), initialize the (θ,ρ) space, N(θ,ρ) = 0; for each pixel point (x,y), find the (θ,ρ) pair that satisfies xcosθ + ysinθ = ρ in the parameter space, and let N(θ,ρ) = N(θ,ρ) + 1; count the magnitudes of all N(θ,ρ), and take out the parameters where N(θ,ρ) > τ. Among them, for the points not less than the threshold, the line is obtained through the inverse mapping of the parameter pair (r,θ) of the Hough space, and its calculation formula is:
[0046]
[0047] In the formula, r is the perpendicular distance from the origin to the line, θ l is the angle between this line and the positive direction of the x-axis, expressed in radians, and (x,y) is the coordinate point in the image space;
[0048] S44. According to the positions of the dashboard center, zero scale line, maximum scale line, and the pointer line, use matplotlib to draw a simplified dashboard schematic diagram; during the drawing process, make a judgment. The judgment process is as follows: Judge the vertical coordinate size relationship between the dashboard center O and the two corner points A and B. If O is less than A or B, then rotate the simplified dashboard schematic diagram by 180 degrees; judge the relative positions of the zero scale and the maximum scale lines. If the line where the zero scale is located is on the right side of the line where the maximum scale is located, then flip the simplified dashboard schematic diagram left and right.
[0049] Preferably, in S5, the specific steps to read the meter according to the simplified dashboard schematic diagram, the meter range, and the measurement unit are as follows:
[0050] S51. Calculate the deflection angle θ1 between the zero scale line and the pointer line and the deflection angle θ2 between the zero scale line and the maximum scale line based on the central angle algorithm;
[0051] Among them, the formula for calculating the deflection angle θ1 between the zero scale line and the maximum scale line is as follows:
[0052]
[0053] In the formula, O(x0,y0) is the center coordinate, A(x1,y1) and B(x2,y2) are the intersection points of the lines where the zero scale line and the maximum scale line are located and the circle respectively, and the vector
[0054]
[0055] Wherein, O(x0, y0) is the center coordinate of the circle, A(x1, y1) and C(x3, y3) are the intersection points of the zero scale line and the line where the pointer is located with the circle respectively, and the vector
[0056] S52. Calculate the ratio of the deflection angle between the zero scale line and the line where the pointer is located to the deflection angle between the zero scale line and the maximum scale line, and then multiply it by the instrument range, combined with the measurement unit, to obtain the reading result. The calculation formula is:
[0057]
[0058] Wherein, K is the instrument range, U is the measurement unit, and result is the reading result.
[0059] Therefore, the present invention proposes an intelligent reading method for industrial pointer-type dial instruments facing complex environments, and its beneficial effects are as follows:
[0060] (1) Based on the deep learning Yolov10 detection framework, the designed dial detection model of the present invention can better identify small targets, thereby improving the reading accuracy; compared with the traditional manual reading method, the reading speed is greatly improved, and the implementation cost is reduced; compared with the existing deep detection framework, a new loss function is designed, bringing higher performance and low inference latency.
[0061] (2) The present invention performs a series of denoising processes on the dial, eliminating the influence of abnormal readings on industrial production in cases where the industrial production environment is complex and changeable, and the dial surface may be inclined, oily, cracked, etc.
[0062] (3) The deep learning model of the present invention can automatically learn the characteristics of the scales in the instrument from a large amount of instrument image data, and can realize the extraction of the scales in the pointer instrument, process the extracted pictures and perform formula calculations. This simple method replaces the cumbersome manual reading procedure and reduces the consumption of human resources.
[0063] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0064] Figure 1 is a schematic flowchart of an intelligent reading method for industrial pointer-type dial instruments facing complex environments of the present invention;
[0065] Figure 2It is the overall flowchart of an intelligent reading method for industrial pointer-type dial instruments facing complex environments according to the present invention;
[0066] Figure 3 It is a deep learning model based on YOLOv10 in an intelligent reading method for industrial pointer-type dial instruments facing complex environments according to the present invention. Specific implementation manners
[0067] To make the technical solutions, advantages and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be described clearly and completely below. The described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the protection scope of this application.
[0068] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention belongs.
[0069] As Figure 1 shown, according to a flowchart of an intelligent reading method for industrial pointer-type dial instruments facing complex environments provided by an embodiment of the present invention, it specifically includes the following steps:
[0070] S1. Detect the instrument panel included in the input image;
[0071] S2. Denoise the instrument panel image;
[0072] S3. Identify the instrument range and measurement unit based on the instrument panel image;
[0073] S4. Identify the center of the instrument panel, zero scale, maximum scale and the scale pointed by the pointer based on the instrument panel image, and construct a simplified schematic diagram of the instrument panel;
[0074] S5. Read the instrument according to the simplified schematic diagram of the instrument panel, the instrument range and the measurement unit.
[0075] In S1, detecting the instrument panel included in the input image is divided into five stages: acquisition, sample production, data input, model optimization and model inference. The specific steps are as follows:
[0076] S11. Acquisition: Arrange a camera near the industrial pointer-type dial instrument to collect instrument pictures;
[0077] S12. Sample production: Manually label the instrument pictures using the labellmg software, where the instrument image categories are as Figure 3 shown;
[0078] S13. Data Input: Input the dataset into the deep learning network model based on YOLOv10 for model training;
[0079] S14. Model Optimization: To improve the performance of YOLOv10, a target detection loss function based on Shape-IoU is designed to obtain the optimal weights; Shape-IoU can calculate the loss by focusing on the shape and scale of the bounding box, making the bounding box regression more accurate. Shape-IoU can be expressed as:
[0080] L Shape-IoU = 1 - IoU + distance shape + 0.5′Ω shape ;
[0081] In the formula, IoU is the intersection over union of the GT box and the anchor box, and distance shape is the distance between the GT box and the anchor box.
[0082] S15. Model Inference: Input the instrument image into the trained model. The model outputs the detection box parameters and categories of the instrument. Call the function to crop out the detection box, intercept the instrument panel contained in the input image, and save it in the corresponding folder.
[0083] In this embodiment, the instrument dials of some parts that need to be read are small, and the real-time requirement is high. Therefore, the improved YOLOv10 is used for dial detection in this embodiment.
[0084] YOLOv10 proposes a consistent double assignment strategy without NMS training. The model uses two prediction heads during training, one using one-to-many assignment and the other using one-to-one assignment. It can utilize the rich supervision signals of one-to-many assignment during training, while using the prediction results of one-to-one assignment during inference, thus achieving efficient inference without NMS. This brings a significant performance improvement and low inference latency.
[0085] YOLOv10 also adopts an efficiency-accuracy-driven model design. After analyzing the influence of classification error and regression error, it is found that the regression head has a greater impact on the performance of YOLOs. Without significantly affecting the performance, the computational cost of the classification head is appropriately reduced.
[0086] In S2, the specific steps for denoising the instrument panel image are to take different operations according to the instrument categories output in the inference stage of the YOLOv10 model: If it is greasy or broken, first perform image restoration and then image correction; if it is normal, directly perform image correction.
[0087] In S2, IPT model is adopted for image inpainting. The model consists of four parts: a head for extracting features from dashboard pictures, an encoder and a decoder for restoring lost information, and a tail for mapping features to the restored image. IPT can make full use of supervised and self-supervised information. The objective function of IPT is defined as:
[0088] L IPT = λ·L contrastive + L supervised ;
[0089] In the formula, L contrastive is the contrast loss, and L supervised is the supervised loss.
[0090] In S2, image correction is based on the STN algorithm to perform image transformation on the dashboard image. The Spatial Transformer Network (STN) converts the dashboard picture into a frontal dashboard picture and consists of three steps: parameter prediction Localisation net, coordinate mapping Gridgenerator, and pixel acquisition Sampler. The specific operations are as follows:
[0091] S21. Localisation net calculates the parameters required for spatial transformation: Input a Feature map: U ∈ R H×W×C , and after several convolutional and fully connected operations, a regression layer is used to obtain the transformation parameter θ. The formula is as follows:
[0092] θ = f loc (U);
[0093] where θ is the transformation parameter, the dimension of θ depends on the specific transformation type selected by the network, U ∈ R H×W×C is the input feature map, and floc(·) is the localisation network, which can adopt any form;
[0094] S22. Grid generator uses the calculated θ to perform corresponding spatial transformation on the Feature map. The formula is as follows:
[0095]
[0096] In the formula, represents the target pixel position on the output feature map, T θ (G i ) represents the applied transformation, G i represents the regular network, A θ represents the transformation matrix, represents the sampling position on the corresponding input feature map;
[0097] S23. The Sampler extracts pixel values from the input feature map (U) according to the positions in the transformed sampling grid T θ (G i ) to generate the output feature map (V); each obtained position in the sampling grid indicates a specific sampling point on the input feature map (U).
[0098] In S3, the specific steps for identifying the instrument range and measurement unit based on the dashboard image are as follows:
[0099] S31. Call the trained text recognition model PaddleOCR, configure the path frame of the dashboard image to be recognized, store the recognition result in result, and extract the pure text part stored in result and store it in text;
[0100] S32. Match text according to the characteristic information of the instrument range and measurement unit, and classify and distinguish the two. The formula is as follows:
[0101]
[0102] In the formula, Ks represents the candidate set of instrument ranges, and Us represents the candidate set of measurement units.
[0103] S33. Select the one with the largest value in the Ks set as the instrument range K, and select the one that conforms to the actual situation in the Us set as the measurement unit U.
[0104] In S4, the specific steps for identifying the center of the dashboard, zero scale, maximum scale, and pointer pointing scale based on the dashboard image and constructing a simplified dashboard schematic diagram are as follows:
[0105] S41. Use the trained semantic segmentation model Deeplabv3+ to segment the scales in the dashboard image, read the segmentation mask. The Deeplabv3+ model consists of two parts: Encoder and Decoder. The main body of the Encoder is a DCNN with dilated convolution, and the classification network Xception is adopted. Under the same computing resources, it can achieve the same accuracy as ResNet and has better expressive ability. The Decoder module further fuses the low-level features and high-level features to improve the accuracy of the segmentation boundary;
[0106] The specific operation steps are as follows:
[0107] After the output layer of DeepLabv3+, the cross-entropy loss is used to calculate the loss between the predicted class map and the true label, and according to the calculated loss, the weights of the network are updated through the backpropagation algorithm. The cross-entropy loss function formula is:
[0108] L = -∑i y i log(p i )-(1 - y i )log(1 - p i );
[0109] In the formula, p is the probability that the model predicts belongs to the instrument scale category, and y is the true label. Among them, y = 1 indicates that the pixel belongs to the instrument scale, and y = 0 indicates that the pixel does not belong to the instrument scale;
[0110] S42. Fit the circle where the scale is located and calculate the center coordinates O(x0, y0). Use Thomas corner detection to obtain the two endpoints A(x1, y1) and B(x2, y2) of the fitting circular curve of the scale. Connect the two endpoints and the center respectively to obtain the zero scale line and the maximum scale line; among them, the center of the instrument panel is the center coordinates. Perform ellipse fitting and Thomas corner detection on the picture grid_image and output the processed picture;
[0111] S43. Determine the position of the line where the pointer is located based on the Hough transform line detection. Select the longest line object as the pointer and store the position information of the pointer; specifically, detect the pointer object based on the Hough transform line detection. Since there are also scale lines and other lines in the instrument panel in addition to the pointer, in most cases, more than one line object is detected, and the length of the pointer is the longest among them. Therefore, select the longest line object as the pointer and store the position information of the pointer.
[0112] The specific steps to determine the position of the line where the pointer is located are as follows:
[0113] Associate each line of the image with the Hough space of a pair of parameters (r, θ), initialize the (θ, ρ) space, N(θ, ρ)=0; for each pixel point (x, y), find the (θ, ρ) pair that satisfies xcosθ + ysinθ = ρ in the parameter space, and let N(θ, ρ)=N(θ, ρ)+1; count the magnitudes of all N(θ, ρ), and take out the parameters where N(θ, ρ)>τ. Among them, for the points not less than the threshold, the line is obtained through the inverse mapping of the parameter pair (r, θ) of the Hough space, and its calculation formula is:
[0114]
[0115] In the formula, r is the perpendicular distance from the origin to the line, θ l is the angle between this line and the positive direction of the x-axis, expressed in radians, and (x, y) is the coordinate point in the image space;
[0116] S44. According to the positions of the center of the dashboard, the zero scale line, the maximum scale line, and the line where the pointer is located, use matplotlib to draw a simplified dashboard schematic diagram; during the drawing process, make judgments. The judgment process is as follows: Judge the vertical coordinate size relationship between the center O of the dashboard and the two corner points A and B. If O is less than A or B, then rotate the simplified dashboard schematic diagram by 180 degrees; judge the relative positions of the zero scale line and the maximum scale line. If the line where the zero scale is located is on the right side of the line where the maximum scale is located, then flip the simplified dashboard schematic diagram horizontally.
[0117] In S5, the specific steps for reading the instrument according to the simplified dashboard schematic diagram, the instrument range, and the unit of measurement are as follows:
[0118] S51. Calculate the deflection angle θ1 between the zero scale line and the line where the pointer is located and the deflection angle θ2 between the zero scale line and the maximum scale line based on the central angle algorithm;
[0119] Among them, the formula for calculating the deflection angle θ1 between the zero scale line and the maximum scale line is as follows:
[0120]
[0121] In the formula, O(x0,y0) is the center coordinate, A(x1,y1) and B(x2,y2) are the intersection points of the lines where the zero scale line and the maximum scale line are located with the circle respectively, and the vector
[0122]
[0123] In the formula, O(x0,y0) is the center coordinate, A(x1,y1) and C(x3,y3) are the intersection points of the lines where the zero scale line and the pointer are located with the circle respectively, and the vector
[0124] S52. Calculate the ratio of the deflection angle between the zero scale line and the line where the pointer is located to the deflection angle between the zero scale line and the maximum scale line, then multiply it by the instrument range, and combine the unit of measurement to obtain the reading result. The calculation formula is:
[0125]
[0126] In the formula, K is the instrument range, U is the unit of measurement, and result is the reading result.
[0127] Therefore, the present invention provides an intelligent reading method for industrial pointer-type dial instruments facing complex environments, which can achieve rapid and high-precision identification of instrument readings, eliminate the impact of abnormal readings on industrial production, and greatly improve the reading speed and accuracy compared with traditional manual reading methods, while reducing the implementation cost.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An intelligent reading method for industrial pointer-type dial instruments facing complex environments, characterized in that, It includes the following steps: S1. Detect the dashboard included in the input image; S2. Denoise the dashboard image; S3. Identify the instrument range and measurement unit based on the dashboard image; S4. Identify the center of the dashboard, zero scale, maximum scale, and the scale pointed by the pointer based on the dashboard image, and construct a simplified dashboard schematic diagram; S5. Read the instrument according to the simplified dashboard schematic diagram, instrument range, and measurement unit.
2. The intelligent reading method for industrial pointer-type disc meters facing complex environments according to claim 1, characterized in that In S1, the detection of the dashboard included in the input image is divided into five stages: acquisition, sample annotation, data input, model optimization, and model inference. The specific steps are as follows: S11. Acquisition: Arrange a camera near the industrial pointer-type dial instrument to collect instrument pictures; S12. Sample annotation: Manually annotate the instrument pictures using the labellmg software; S13. Data input: Input the data set into the deep learning network model based on YOLOv10 for training the model; S14. Model optimization: Optimize through the loss function to obtain the best weights; S15. Model inference: Input the instrument image into the trained model. The model outputs the detection box parameters and category of the instrument. Call the function to crop out the detection box, intercept the dashboard included in the input image, and store it in the corresponding folder.
3. The intelligent reading method for industrial pointer-type dial instruments facing complex environments according to claim 2, characterized in that In S14, the loss function is Shape-IoU, which can be expressed as: L Shape-IoU = 1 - IoU + distance shape + 0.5 × Ω shape ; where IoU is the intersection over union of the GT box and the anchor box, and distance shape is the distance between the GT box and the anchor box 4. The intelligent reading method of industrial pointer-type dial instruments for complex environments according to claim 1, characterized in that, In S2, the specific steps for denoising the dashboard image are to take different operations according to the instrument category output in the inference stage of the YOLOv10 model: If it is greasy or broken, first perform image restoration and then image correction; if it is normal, directly perform image correction.
5. The intelligent reading method for industrial pointer-type dial instruments facing complex environments according to claim 4, characterized in that, In S2, IPT model is used for image restoration. The model consists of four parts: a head that extracts features from the dashboard picture, an encoder and a decoder that restore the lost information, and a tail that maps the features to the restored image. The objective function of IPT is defined as: L IPT = λ·L contrastive + L supervised ; where L contrastive is the contrastive loss, and L supervised is the supervision loss.
6. The intelligent reading method for industrial pointer-type dial instruments facing complex environments according to claim 4, characterized in that In S2, image correction is based on the STN algorithm to perform image transformation on the dashboard image; The spatial transformation network STN converts the dashboard picture into a positive dashboard picture, and it includes three steps: parameter prediction Localisation net, coordinate mapping Gridgenerator, and pixel acquisition Sampler. The specific operations are as follows: S21. The localisation net calculates the parameters required for spatial transformation: Input a Feature map: U ∈ R H ×W×C , after several convolutional and fully connected operations, a regression layer is borrowed to obtain the transformation parameter θ, and the formula is as follows: θ = f loc (U); where θ is the transformation parameter, and the dimension of θ depends on the specific transformation type selected by the network. U ∈ R H×W×C is the input feature map, floc(·) is the localization network, and the localization network floc(·) can take any form; S22. Grid generator performs corresponding spatial transformation on the Feature map using the calculated θ. The formula is as follows: In the formula, represents the target pixel position on the output feature map, T θ (G i ) represents the applied transformation, G i represents the regular network, A θ represents the transformation matrix, represents the sampling position on the corresponding input feature map; S23. The Sampler extracts pixel values from the input feature map (U) according to the positions in the transformed sampling grid T θ (G i ) to generate the output feature map (V); each obtained position in the sampling grid indicates a specific sampling point corresponding to the input feature map (U).
7. The intelligent reading method for industrial pointer-type disk instruments in complex environments according to claim 1, wherein In S3, the specific steps for identifying the instrument range and measurement unit based on the dashboard image are as follows: S31. Call the trained text recognition model PaddleOCR, configure the path frame of the dashboard picture to be recognized, store the recognition result in result, and extract the pure text part stored in result and store it in text; S32. Match text according to the characteristic information of the instrument range and measurement unit, and classify and distinguish the two. The formula is as follows: In the formula, Ks represents the candidate set of instrument ranges, and Us represents the candidate set of measurement units. S33. Select the one with the largest value in the Ks set as the instrument range K, and select the one that matches the actual situation in the Us set as the measurement unit U.
8. The intelligent reading method for industrial pointer-type disk instruments facing complex environments according to claim 1, characterized in that In S4, the specific steps to construct a simplified instrument panel schematic diagram based on the recognition of the instrument panel center, zero scale, maximum scale, and pointer pointing scale in the instrument panel image are as follows: S41. Use the trained semantic segmentation model Deeplabv3+ to segment the scales in the instrument panel image and read the segmentation mask. The specific operation steps are as follows: After the output layer of DeepLabv3+, the cross-entropy loss is used to calculate the loss between the predicted category map and the true label, and according to the calculated loss, the weights of the network are updated through the backpropagation algorithm. The cross-entropy loss function formula is: L = -∑ i y i log(p i ) - (1 - y i ) log(1 - p i ); In the formula, p is the probability predicted by the model belonging to the instrument scale category, and y is the true label. Among them, y = 1 indicates that the pixel belongs to the instrument scale, and y = 0 indicates that the pixel does not belong to the instrument scale; S42. Fit the circle where the scale is located and calculate the center coordinates O(x0, y0). Use the Thomas corner detection to obtain the two endpoints A(x1, y1) and B(x2, y2) of the fitting circular curve of the scale, and connect the two endpoints and the center respectively to obtain the zero scale line and the maximum scale line; among them, the instrument panel center is the center coordinates. Perform ellipse fitting and Thomas corner detection on the picture grid_image and output the processed picture; S43. Determine the position of the line where the pointer is located based on the Hough transform line detection. Among them, select the longest line object as the pointer and store the position information of the pointer; the specific steps to determine the position of the line where the pointer is located are: Associate each line in the image with a pair of parameters (r, θ) in the Hough space, initialize the (θ, ρ) space, and N(θ, ρ) = 0; for each pixel point (x, y), find the (θ, ρ) pair that satisfies xcosθ + ysinθ = ρ in the parameter space, and let N(θ, ρ) = N(θ, ρ) + 1; count the magnitudes of all N(θ, ρ), and take out the parameters where N(θ, ρ) > τ. Among them, for the points not less than the threshold, the line is obtained through the inverse mapping of the parameter pair (r, θ) of the Hough space, and its calculation formula is: where r is the perpendicular distance from the origin to the line, θ l is the angle between this line and the positive x-axis, expressed in radians, and (x, y) is the coordinate point in the image space; S44. According to the positions of the instrument panel center, zero scale line, maximum scale line, and the line where the pointer is located, use matplotlib to draw a simplified instrument panel schematic diagram; during the drawing process, make a judgment. The judgment process is: judge the vertical coordinate size relationship between the instrument panel center O and the two corner points A and B. If O is less than A or B, rotate the simplified instrument panel schematic diagram by 180 degrees; judge the relative positions of the zero scale and the maximum scale lines. If the line where the zero scale is located is on the right side of the line where the maximum scale is located, flip the simplified instrument panel schematic diagram left and right.
9. The intelligent reading method for industrial pointer-type disk instruments facing complex environments according to claim 1, wherein, In S5, the specific steps to read the instrument according to the simplified instrument panel schematic diagram, instrument range, and measurement unit are as follows: S51. Calculate the deflection angle θ1 between the zero scale line and the line where the pointer is located and the deflection angle θ2 between the zero scale line and the maximum scale line based on the central angle algorithm; Among them, the formula for calculating the deflection angle θ1 between the zero scale line and the maximum scale line is as follows: Where O(x0, y0) is the center coordinate of the circle, A(x1, y1) and B(x2, y2) are the intersection points of the lines where the zero scale line and the maximum scale line are located with the circle respectively, and the vector Wherein, O(x0, y0) is the center coordinate of the circle, A(x1, y1) and C(x3, y3) are the intersection points of the zero scale line and the line where the pointer is located with the circle respectively, and the vector S52. Calculate the ratio of the deflection angle between the zero scale line and the straight line where the pointer is located to the deflection angle between the zero scale line and the maximum scale line, then multiply by the instrument range, and combine with the measurement unit to obtain the reading result. The calculation formula is: In the formula, K is the instrument range, U is the measurement unit, and result is the reading result.
Citation Information
Patent Citations
Automatic identification reading method and system for pointer instrument
CN112818988A
Pointer instrument automatic reading method based on improved semantic segmentation network
CN114266881A