A Method for Automatic Reading of Circular Pointer Instruments Based on Deep Learning
Through the modular deep learning method, the versatility and adaptability of the instrument automatic reading recognition technology are solved, and efficient automatic reading recognition in complex environments is achieved, which is suitable for various circular pointer instruments.
Patent Information
- Application Number
- CN202211144633.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-09-20
AI Technical Summary
The existing automatic instrument reading recognition technology lacks high versatility and relies on prior information and scenario limitations, so it is impossible to effectively identify pointer instrument readings in complex environments.
Using a deep learning-based method, the automatic recognition of instrument readings is realized through modular steps such as instance segmentation, image enhancement, dial tilt correction, dial character and pointer information extraction, including instance segmentation deep learning model, super-resolution reconstruction and bilateral filtering and other technologies to process images, and calculate readings in combination with the angle method.
It realizes efficient automatic reading recognition of circular pointer instruments without relying on prior information and scene limitations, and has strong versatility and adaptability. It can quickly apply cutting-edge algorithms in different fields, enhancing the practicality and accuracy of the method.
Smart Images

Figure CN115546795B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of instrument reading, and in particular to an automatic reading method for circular pointer-type instruments based on deep learning. Background Art
[0002] Instruments are the general term for instruments that display numerical values. They are important tools for measuring various data in production and life and play an important role in understanding the environmental state. Although the digital age has arrived and a large amount of sensor data can be directly transmitted to a computer, in many traditional production scenarios such as substations, due to complex electromagnetic environments or un-updated equipment, pointer-type instruments still play an important role in production and life.
[0003] At present, the main method for obtaining data of pointer-type instruments is manual transcription, that is, the production unit arranges special staff for inspection, and then reads each instrument for data verification and recording. This method has a dull and single job content, and in some operation scenarios with a long inspection route, it also has relatively high physical requirements for the inspection personnel. With the development of technology, artificial intelligence technology has become a trend of the times, and using inspection robots to replace workers for inspection has attracted more and more researchers' attention.
[0004] The automatic reading and recognition technology of instruments is one of the core technologies for the instrument reading inspection task of inspection robots. The current automatic reading and recognition technology of instruments mainly relies on digital image processing technology and pattern recognition technology, and can be mainly divided into methods based on traditional image processing, methods based on template matching, and methods based on deep learning. The main difference between these three methods lies in the different methods of obtaining the ROI image of the dial. The traditional image processing method mainly uses the Hough circle detection method to obtain the instrument ROI; the method based on template matching mainly uses the instrument template to obtain the corresponding instrument dial area from the scene image; the method based on deep learning mainly uses the object detection or instance segmentation model to obtain the rectangular bounding box or segmentation mask of the instrument dial area from the scene. After obtaining the instrument dial image, use line detection algorithms such as the Hough transform to obtain the fitted line of the pointer, and then based on the angle method or the distance method, realize the reading calculation. Such methods usually regard the scale information of the instrument as prior information and only focus on the extraction of the dial pointer information. The methods for extracting dial information are mostly traditional image processing methods, which are sensitive to the parameters of the algorithm. Once the shooting light and shadow conditions change or the type of instrument changes, the algorithm often fails. Due to the complexity of the instrument problem that its reading is not presented in the image but hidden in the relationship between the pointer and the scale points, various deep learning models in the field of computer vision currently cannot learn such complex representations, so an end-to-end instrument reading recognition system cannot be realized through deep learning.
[0005] The various deficiencies of the above existing algorithms have led to the lack of a highly general method in the current instrument recognition algorithm, with insufficient intelligence. In engineering applications, it is often necessary to adjust parameters on-site or impose various restrictions on the application scenarios. Therefore, designing a method for automatic reading of instruments that is as general as possible remains a challenging problem to be solved. Summary of the Invention
[0006] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose an automatic reading method for circular pointer-type instruments based on deep learning. The instrument reading problem is decomposed into sub-problems of several modules, and then different computer vision deep learning models are used to solve the corresponding sub-problems respectively, and finally the reading result of the instrument is obtained. The present invention has good generality and accuracy, does not rely on the prior information of the instrument to be detected, and does not require restrictions on the instrument scene, providing a general and effective solution for solving the problem of intelligent reading of instruments.
[0007] To achieve the above purpose, the technical solution provided by the present invention is: an automatic reading method for circular pointer-type instruments based on deep learning, including the following steps:
[0008] 1) Obtain an image of the inspection site through a visible light camera, which is called a scene image;
[0009] 2) Use a dial area extraction module developed based on deep learning instance segmentation technology to detect the instrument dial area in the scene image, segment and extract the instrument dial image from the scene image, remove the background part, and obtain the part that only contains the instrument dial, which is called the dial ROI image;
[0010] 3) By comprehensively considering the size and blur degree of the dial ROI image, obtain an index for measuring the readability of the instrument, which is called the instrument readability index. According to the calculation result of this index, match the corresponding image enhancement module to perform image enhancement on the dial ROI image;
[0011] 4) Use a dial tilt correction module developed based on image transformation technology to solve the dial tilt problem caused by the camera not being directly facing the plane where the instrument is located during the shooting process. The dial image after image enhancement and dial correction is called a high-quality dial ROI image;
[0012] 5) Use deep learning technology to extract the instrument dial information. Specifically, use a dial character information extraction module developed based on deep learning OCR technology to obtain the character information of the high-quality dial ROI image and a dial pointer information extraction module developed based on deep learning key point detection technology to obtain the pointer information of the high-quality dial ROI image;
[0013] 6) Use a pointer scale interval matching module to obtain the effective information of the instrument reading from the extraction result of the above instrument dial information. The effective information of the instrument reading respectively refers to: the end point of the pointer, the rotation center point of the pointer, and the upper and lower bound scale numbers of the scale interval where the pointer is located;
[0014] 7) Use an instrument reading module to calculate the instrument reading based on the above effective information of the instrument reading.
[0015] Furthermore, in step 2), the dial area extraction module is a trained instance segmentation deep learning model. Through this model, the rectangular bounding box coordinates and image mask of the instrument dial in the scene image can be obtained. Then, the sub-image within the range of the above rectangular bounding box is segmented from the scene image, and the part outside the dial area in the above sub-image is removed using the image mask to obtain the part that only contains the instrument dial, which is called the dial ROI image.
[0016] Furthermore, in step 3), the instrument readability index is denoted as F, and the calculation formula of this index is as follows:
[0017]
[0018] In the formula, var represents the image variance, which describes the degree of blurriness of the image; H and W respectively represent the height and width of the image, which describe the size of the image; α and β are two weight coefficients used to adjust the influence weights of the degree of blurriness and size on the calculation result of the instrument readability index;
[0019] Among them, the image variance var represents the calculation result of summing the squares of the differences between the gray values of all pixels and the average gray value of the image, and then dividing by the total number of pixels. The specific calculation process includes the following steps:
[0020] 3.1) Convert the color dial ROI image into a dial ROI grayscale image;
[0021] 3.2) Perform convolution operation on the dial ROI grayscale image using the following Laplacian operator template:
[0022]
[0023] 3.3) Calculate the variance σ of the dial ROI grayscale image after convolution using the following formula 2 :
[0024]
[0025] In the formula, x represents the pixel value of a point on the image, μ represents the gray mean value of the image, the summation range of the numerator includes all pixels in the image, and W*H represents the total number of pixels in the image.
[0026] Further, in step 3), the image enhancement module includes a super-resolution reconstruction deep learning model and a bilateral filtering module. According to the calculation results of different meter readability indices, different image enhancement methods are selected to enhance the dial ROI image.
[0027] According to the meter readability index F of the input dial ROI image, if F is greater than a preset threshold, it is determined that the dial ROI image is a "low-readability image", and the super-resolution reconstruction deep learning model is used to perform image super-resolution reconstruction processing on the input image; if F is less than the preset threshold, it is determined that the dial ROI image is a "non-low-readability image", and the bilateral filtering module is used to perform conventional image enhancement processing on the image.
[0028] The calculation formula of the bilateral filtering module is:
[0029]
[0030] In the formula, f(x, y) is the response at the position of the pixel point (x, y) after filtering, g(x, y) is the value of each pixel point in the neighborhood of the pixel point (x, y), and W(x, y) is the combined weight coefficient of each pixel point; the function of this bilateral filtering module is to perform smoothing filtering on the rest while keeping the edges clear to remove image noise and improve image quality.
[0031] Further, in step 4), the dial tilt correction module specifically performs the following operations:
[0032] 4.1) Obtain the meter image mask obtained by the dial area extraction module, which is a binary mask image of the meter contour, mask.
[0033] 4.2) Perform ellipse fitting on the above mask image to obtain the contour fitting ellipse parameters of the dial ROI image, including the ellipse center point, the length of the major axis of the ellipse, the length of the minor axis of the ellipse, and the angle between the major axis of the ellipse and the vertical direction.
[0034] 4.3) Obtain the coordinates of the endpoints of the major axis and minor axis of the ellipse obtained after fitting, and calculate the coordinates of the point on the straight line where the minor axis is located and the distance to the ellipse center point is the length of the major axis, which is called the correction expected coordinates of the minor axis of the ellipse.
[0035] 4.4) Use the coordinates of the major axis and the endpoints of the minor axis of the ellipse, the coordinates of the endpoints of the major axis of the ellipse, and the correction expected coordinates of the minor axis of the ellipse to form 4 groups of feature point pairs, and use these feature point pairs to calculate the projective transformation matrix. The calculation of the projective transformation matrix is as follows:
[0036] p2 = H' * p1
[0037] In the formula, p1 represents the original before transformation Figure 1Let p1 be a point, p2 be the corresponding feature point of p1, and H′ be the projective transformation matrix, which represents the process of mapping point p1 in the original image to point p2 in the transformed image. Expanding the above formula gives:
[0038]
[0039] In the formula, (x1, y1) represents the coordinates of point p1, (x2, y2) represents the coordinates of point p2, and H 11 ~H 33 are all parameters of matrix H′. From the above matrix form, a system of equations can be obtained. Among them, from H 11 ~H 33 These 9 parameters can be used to construct a system of equations with the above 4 sets of feature point pairs to solve for the unique projective transformation matrix H′;
[0040] 4.5) The projective transformation matrix H′ obtained by using the formula in step 4.4) can globally process the dial ROI image, calculate the pixel coordinates of each dial ROI image after correction before correction, so as to realize the correction of the elliptical dial ROI image into a circular shape.
[0041] Furthermore, in step 5), the dial character information to be extracted by the dial character information extraction module includes the character text content of each character image and its corresponding position coordinates. The acquisition process specifically includes the following steps:
[0042] 5.1.1) Input the high-quality dial ROI image after image enhancement and dial correction into a trained deep learning model for scene text detection. The model infers and outputs all character region information of the high-quality dial ROI image. The character region information is the rectangular bounding box of the character, which is described by the four vertex coordinates of the rectangular bounding box;
[0043] 5.1.2) Cut all characters from the high-quality dial ROI image in sequence through the rectangular bounding box coordinates of the character image to obtain a sequence of character images of the high-quality dial ROI image;
[0044] 5.1.3) Use the vertex coordinates of the rectangular bounding box of each character image to calculate the character position coordinates. The character position coordinates are defined as the center coordinates of the vertices of the character rectangular bounding box, and their horizontal and vertical coordinates are the averages of the four vertex coordinates of the rectangular bounding box;
[0045] 5.1.4) Send each character image in the sequence of character images into a trained deep learning model for text recognition of character images in turn to obtain the text content recognition results of each character image in the sequence of character images, and save the text content and the corresponding position coordinates in pairs, that is, the dial character information is obtained.
[0046] Further, in step 5), the dial pointer information extraction module specifically performs the following operations:
[0047] 5.2.1) Detect the pointers inside the dial through a target detection deep learning model to obtain the rectangular bounding box parameters of the pointers, including the center point position coordinates of the pointers and the width and height of the rectangular bounding box;
[0048] 5.2.2) Based on the rectangular bounding box parameters of the pointers, obtain the sub-image within the range of the pointer rectangular bounding box from the scene image, and the obtained image is called the pointer ROI image;
[0049] 5.2.3) Input the pointer ROI image into a trained key point detection deep learning model, and directly infer the pointer rotation center point and the position coordinates of the pointer end point through this model. The above two points and their positions are the extracted pointer information.
[0050] Further, in step 6), the pointer scale interval matching module is divided into two stages: scale number screening and matching the scale interval where the pointer is located:
[0051] Stage 1: Scale number screening:
[0052] 6.1.1) Define the scale numbers as the digital characters representing the instrument scales in the dial ROI image. First, screen out all the characters with pure digital text content according to the character text content obtained by the dial information extraction module as the preliminary screening result of the scale numbers, hereinafter referred to as digital characters;
[0053] 6.1.2) Calculate the Euclidean distance from the above-screened digital characters to the pointer rotation center point to obtain a distance sequence;
[0054] 6.1.3) Use the k-means clustering algorithm to cluster the above distance sequence. Set the number of clustering categories k to 3. Since the scale numbers of the pointer-type instrument are approximately distributed on the same circle in space, the distances to the rotation center point are approximately the same and are clustered into the same category. The digital characters within this category are used as the effective screening result of the digital characters, and the remaining digital characters are determined as invalid screening results and removed from the digital characters;
[0055] 6.1.4) Calculate the rotation reference angle of each remaining digital character. The rotation reference angle is defined as the angle between the line connecting the digital character and the pointer rotation center point and the ray vertically downward from the pointer rotation center point with the line passing through the rotation center point downward as the reference line;
[0056] 6.1.5) Use the digital character text content and the corresponding rotation reference angle to form a <digital character text content, rotation reference angle> key-value pair, and the key-value pairs of all the remaining digital characters form a key-value pair sequence;
[0057] 6.1.6) Sort the above key-value pair sequence in ascending order using the digital character text content as the sorting keyword, and check whether the rotation reference angles of each key-value pair in the sorted key-value pair sequence also conform to the ascending order rule. If there are key-value pairs whose rotation reference angles do not conform to the ascending order rule, the corresponding digital characters of these key-value pairs are determined as non-scale digital characters and removed from the digital characters;
[0058] 6.1.7) The characters remaining after the above screening are the final scale numbers;
[0059] Stage two, match the scale interval where the pointer is located:
[0060] 6.2.1) Calculate the Euclidean distance from the end point of the pointer to each scale number, and use the two scale numbers closest to the end point of the pointer as the upper and lower bound scale numbers of the scale interval;
[0061] 6.2.2) Define the angle between two points as the angle of the acute angle formed by the lines connecting the two points to the center of rotation of the pointer with the center of rotation of the pointer as the vertex; according to the above definition, the angles between the end point of the pointer and the upper and lower bound scale numbers of the to-be-determined scale interval and the angle between the upper and lower bound scale numbers of the to-be-determined scale interval can be calculated;
[0062] 6.2.3) If the sum of the angles from the end point to the upper and lower bound scale numbers of the to-be-determined scale interval is approximately equal to the angle between the upper and lower bound scale numbers of the to-be-determined scale interval, the match is successful; if not, introduce the digital character closest to the end point of the pointer among the remaining digital characters to form a to-be-determined scale interval with the above two scale numbers respectively, and repeat the above matching process until the match is successful. Among the successfully matched digital characters, the character with the larger digital character text content is called the upper bound number of the scale interval, and the character with the smaller digital character text content is called the lower bound number of the scale interval.
[0063] Furthermore, in step 6.1.4), for the angle calculation of any three non-collinear point coordinates, specifically as follows:
[0064] Use the Pythagorean theorem formula to calculate the length of the line segment between two points:
[0065]
[0066] In the formula, a′ represents the length of the line segment, (x b , y b ) and (x c , y c ) are the endpoint coordinates of both ends of the line segment respectively. Based on the above formula, the lengths of the three sides of the triangle formed by any three non-collinear points can be calculated;
[0067] Calculate the angle from the side lengths using any angle calculation formula for solving triangles:
[0068]
[0069] In the formula, a is the length of the opposite side segment of the vertex where the angle to be calculated is located, b and c are the lengths of the adjacent side segments of the vertex where the angle to be calculated is located respectively, and A1 is the angle of the angle to be calculated.
[0070] Furthermore, in step 7), the instrument reading module is improved from the angle method, and the formula is expressed as follows:
[0071]
[0072] In the formula, Reading represents the finally obtained instrument reading; max_scale and min_scale respectively represent the digital character text contents of the upper bound number and the lower bound number of the scale interval, which can be directly obtained after matching and determining the upper and lower bound scale numbers; ang_end represents the angle between the center point of the pointer rotation and the lower bound number of the scale interval; ang_interval represents the angle between the upper bound number and the lower bound number of the scale interval. When the position coordinates of the center point of the pointer rotation, the end point of the pointer, the upper bound number and the lower bound number of the scale interval are determined, the angle values of ang_end and ang_interval can be calculated, and the calculation method is the same as the angle calculation method for any three non-collinear point coordinates in step 6.1.4).
[0073] The result calculated by the above formula is the final result of the automatic reading recognition of the circular pointer type instrument.
[0074] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0075] 1. The method of the present invention has strong versatility and can automatically recognize the readings of all circular instruments of single reading systems without relying on prior knowledge, that is, regardless of the scene and the specific instrument model, and its versatility is significantly better than the existing mainstream methods.
[0076] 2. The method of the present invention has the characteristics of modularization. The methods between each stage are relatively independent and can be optimized separately without affecting the implementation of the overall method. In this way, when a more optimal deep learning model appears in different fields, the industry-leading algorithms can be quickly applied to the system to maintain the leading level of performance.
[0077] 3. The method of the present invention designs an image enhancement module with a certain adaptability, which can directly process the problems of instrument blurring and too low instrument resolution in the scene image, can overcome the interference factors brought by the engineering site to a certain extent, and enhances the practicability of the method.
[0078] 4. When analyzing the information on the instrument dial, the method of the present invention introduces a module for extracting the character information of the dial based on OCR technology. Based on this module, relevant information about the instrument type can be further obtained, so as to realize functions such as instrument type recognition and automatic recognition of the unit of the instrument reading, and has better potential for further development. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is the overall flow chart of the method of the present invention.
[0080] Figure 2 is the flow chart for extracting the ROI image of the high-quality dial.
[0081] Figure 3 is the schematic diagram for assisting the dial correction.
[0082] Figure 4 is the flow chart for extracting the effective information of the dial.
[0083] Figure 5 is the schematic diagram of an example of the rotation reference angle.
[0084] Figure 6 is the schematic diagram for assisting the calculation formula of the triangle.
[0085] Figure 7 is the schematic diagram of an example of the instrument reading calculation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0087] As Figure 1 shown, the inspection robot obtains the image of the scene to be detected (i.e., the image of the inspection site) through the visible light camera and inputs it into the dial area extraction module to obtain the ROI image of the dial. Then, according to the calculation result of the instrument readability index of the ROI image of the dial, the corresponding image enhancement module is determined to be used. After the ROI image of the dial is enhanced by the image, the dial is corrected to obtain the high-quality ROI image of the dial. The high-quality ROI image of the dial is input into the module for extracting the character information of the dial and the module for extracting the pointer information of the dial respectively to obtain the effective information of the characters and pointers of the instrument dial. After passing through the pointer scale interval matching module to find the scale interval where the pointer is located, the final instrument reading is calculated by using the instrument reading module. The above modules will be further introduced in detail below in conjunction with more detailed flow charts and schematic diagrams.
[0088] As Figure 2As shown, the input scene image is first fed into a dial area extraction module developed based on deep learning instance segmentation technology. The dial area extraction module is a trained instance segmentation deep learning model, specifically the Mask RCNN model. Through this model, the rectangular bounding box coordinates of the instrument dial in the scene image and the instrument image mask can be obtained, the sub-image within the range of the above rectangular bounding box can be segmented from the scene image, and the part outside the dial area in the above sub-image can be removed using the corresponding instrument image mask to obtain the part that only contains the instrument dial, which is called the dial ROI image.
[0089] Next, the Laplace operator is used to calculate the image variance of the dial ROI image as the evaluation basis for the image blurriness. The calculation process of the above image variance is as follows:
[0090] 1) Convert the color dial ROI image into a dial ROI grayscale image;
[0091] 2) Use the following Laplace operator template to perform convolution operation on the dial ROI grayscale image:
[0092]
[0093] 3) Use the following formula to calculate the variance var of the dial ROI grayscale image after convolution (i.e., the image variance of the dial ROI image), and use this index as the judgment index for blurriness:
[0094]
[0095] In the formula, x represents the pixel value of a point on the image, μ represents the grayscale mean of the image, the summation range of the numerator part on the right side of the formula includes all pixels in the image, H and W respectively represent the height and width of the image, and W*H represents the total number of pixels in the image.
[0096] Combined with the image variance of the dial ROI image and the size of the dial ROI image, a metric index called "instrument readability index" is proposed, denoted as F, and its calculation formula is as follows:
[0097]
[0098] In the formula, var represents the obtained image variance; α and β are two weight coefficients used to adjust the weights of blurriness and size in the calculation of the instrument readability index. β is set as the average size of the images in the instrument dial image dataset, and α is set to 1 by default. During subsequent testing, these two weight parameters can be adjusted according to specific needs in engineering practice to obtain the optimal effect.
[0099] The image enhancement module includes a super-resolution reconstruction deep learning model and a bilateral filtering module. According to the calculation results of different meter readability indexes, different image enhancement methods are selected to enhance the dial ROI image.
[0100] When the calculation result of the meter readability index is greater than the set threshold, the image is determined as a "low readability image", and further the dial ROI image is input into the super-resolution reconstruction deep learning model (RealSR model), and super-resolution reconstruction is used to enhance the highly blurred image. On the contrary, if the meter readability index of the above dial ROI image is determined as a "non-low readability image", the image is input into the bilateral filtering module for conventional image enhancement processing. The calculation formula of the above bilateral filtering module is:
[0101]
[0102] In the formula, f(x, y) is the response at the position of the pixel point (x, y) after filtering, g(x, y) is the value of each pixel point in the neighborhood of the pixel point (x, y), and W(x, y) is the combined weight coefficient of each pixel point. The main function of this bilateral filtering module is to perform smoothing filtering on the rest while keeping the edges clear to remove image noise and improve image quality.
[0103] After image enhancement, the meter image mask obtained by the instance segmentation deep learning model is used for image correction to solve the distortion problem caused by the fact that the inspection robot's shooting angle is not directly facing the meter dial. As Figure 3 shown in the schematic diagram for dial correction assistance. What is obtained by shooting at an inclined angle is an approximately elliptical meter contour. After contour detection and ellipse fitting on this meter image mask, an ellipse ABCD as shown by the solid line in Figure 3 can be obtained. In the current mainstream computer vision algorithm library, the fitted ellipse is described by the ellipse center point, the length of the major axis of the ellipse, the length of the minor axis of the ellipse, and the angle between the major axis and the vertical direction. Based on the above information and geometric knowledge, the position coordinates of the four long and short axis intersection points A, B, C, and D of the ellipse can be obtained. Then, the position coordinates of two points B' and D' on the straight line where the minor axis of the meter is located and at a distance from the center of the circle equal to the length of the major axis can be obtained by geometric knowledge. The positions of points A and C remain unchanged and are re-recorded as points A' and C'. A circle A'B'C'D' is obtained as the expected dial contour after correction (as shown by the dotted circle in Figure 3 ), and the position coordinates of the four points A'B'C'D' are the expected correction coordinates of the ellipse.
[0104] Using the 4 groups of feature point pairs composed of ABCD and A'B'C'D', the projective transformation matrix can be obtained, and then the correction of the meter can be completed using the projective transformation. The calculation formula of the projective transformation matrix is as follows:
[0105] p2=H′*p1
[0106] In the formula, p1 represents the original Figure 1 Points, p2 represents the feature point corresponding to p1, H′ is the projection transformation matrix, which represents the process of mapping the point p1 in the original image to the point p2 in the transformed image. Expanding the above formula, we get:
[0107]
[0108] In the formula, (x1, y1) represents the coordinates of point p1, (x2, y2) represents the coordinates of point p2, and H 11 ~H 33 are all parameters of the matrix H′. From the above matrix form, we can get a set of equations; 11 ~H 33 These 9 parameters can be used to construct four sets of equations using the above four sets of characteristic point pairs AA', BB', CC', and DD', and the unique projective transformation matrix H' can be solved.
[0109] The projection transformation matrix can be used to achieve the projection transformation of the entire image, and finally the dial correction task is achieved. After correction, a relatively clear dial ROI image that is directly facing the camera's perspective is obtained, which is called a high-quality dial ROI image.
[0110] like Figure 4 The figure shows the flow chart of effective dial information extraction. The left and right branches in the figure correspond to the specific processes of the dial character information extraction module and the dial pointer information extraction module respectively. After the branches are merged, they enter the pointer scale interval matching module.
[0111] The dial character information extraction module mainly includes two parts: text detection and text recognition. The high-quality dial ROI image is first input into a trained deep learning model for scene text detection (DB model). This model outputs all character region information within the high-quality dial ROI image in the form of the four-point coordinates of the rectangular bounding box. The character region information is the rectangular bounding box of the character, described by the four vertex coordinates of the rectangular bounding box. The sub-images within the range of the rectangular bounding box are segmented to obtain a sequence of character images of the high-quality dial ROI. Using the vertex coordinates of the rectangular bounding boxes of each character image in the sequence of character images, the position coordinates of each character are calculated. The character position coordinates are defined as the center coordinates of the vertices of the character rectangular bounding box, and their horizontal and vertical coordinates are respectively the average values of the four vertex coordinates of the rectangular bounding box. Then, the character images in these sequences of character images are sequentially fed into a trained deep learning model for character image text recognition (CRNN model) to obtain the text recognition results of each character image, that is, the character text content corresponding to the character image. The recognized character text content and the corresponding position coordinates are saved in pairs using data structures such as dictionaries. In this way, the extraction of all character information within the high-quality dial ROI image is completed. In the following text, "character" refers to both the character image and its recognition result, including two attributes: the text content of the character and the position coordinates.
[0112] The dial pointer information extraction module mainly includes a deep learning model for object detection (FasterRCNN model) and a deep learning model for key point detection (HRNet model). The high-quality dial ROI image is first input into the trained deep learning model for object detection to obtain the parameters of the rectangular bounding box of the pointer. The sub-image within the range of the rectangular bounding box is segmented to obtain the pointer ROI image and input into the deep learning model for key point detection. The deep learning model for key point detection directly predicts and infers the pointer key points and the corresponding coordinates. The pointer key points include the pointer rotation center point and the pointer end point. The two key points and the corresponding coordinates are the required effective pointer information, and the line connecting the two key points reflects the line where the pointer is located.
[0113] The pointer scale interval matching module mainly includes two stages: scale number screening and matching the interval where the pointer scale is located.
[0114] In the scale number screening stage, the scale numbers are mainly screened based on the characters obtained from the above-mentioned dial character information extraction module and the pointer rotation center point coordinates obtained from the dial pointer information extraction module. The main purpose is to separate the scale numbers from the characters, and at the same time, avoid the situation where other characters are misjudged as scale numbers or non-scale digital characters are screened as scale numbers. The entire screening process is carried out in three rounds.
[0115] In the first round of screening, based on the character text content recognized from the character image text, all characters with pure numeric character text content are selected as the preliminary screening results of the scale numbers, hereinafter simply referred to as numeric characters.
[0116] In the second round of screening, based on the characteristic that each scale number of the pointer-type instrument is generally distributed on a circle centered on the pointer rotation center, the distance between each numeric character and the pointer rotation center point is used to remove numeric characters with abnormal spatial distribution.
[0117] First, calculate the Euclidean distance between each numeric character and the pointer rotation center point. The calculation formula is as follows:
[0118]
[0119] After the calculation, a distance sequence of each numeric character is obtained.
[0120] Use the k-means algorithm to cluster the distances in the above numeric character distance sequence, and cluster the above distances into three clusters. The cluster with the largest number of numeric characters after clustering is used as the effective screening result of the numeric characters. The other two clusters represent the numeric characters that are too close to and too far from the pointer rotation center point, and are all outliers in the distance sequence, which are determined as invalid screening results and removed from the numeric characters.
[0121] In the third round of screening, based on the characteristic that each scale number of the pointer-type instrument continuously increases with the pointer rotation direction, calculate the reference angle corresponding to each numeric character, and sort it in combination with the scale number text content in the pointer rotation direction, and remove the numeric characters that do not conform to the above rules.
[0122] From Figure 5 As shown in the schematic diagram of the rotation reference angle example, select the ray vertically downward passing through the rotation center point as the reference line, and define the angle between the line connecting the numeric character X and the pointer rotation center point O and the reference line as the rotation reference angle corresponding to the numeric character X. The reference line is a ray starting from the pointer rotation center point and vertically downward. In the specific calculation process, any point on the ray below point O can be selected as the reference point Y, as long as the abscissa of Y is the same as that of point O and the ordinate is below point O. Use the coordinates of the three points O, X, and Y to calculate the angle value with point O as the vertex, which is the rotation reference angle.
[0123] Refer to Figure 6 As shown in the schematic diagram of the triangle calculation formula for assistance, the method for calculating the angle value based on three points is as follows:
[0124] Using the Pythagorean theorem formula, the line segment length between any two points can be calculated. Taking the calculation of the length of side a in Figure 6 as an example for illustration, the calculation formula is as follows:
[0125]
[0126] In the formula, a represents the length of the side, B1(x b , y b ) and C1(x c , y c ) are the endpoint coordinates of both ends of the line segment a respectively. Based on the above formula, the lengths of the other two sides in Figure 6 can be calculated, and are denoted as b and c respectively;
[0127] Using any angle calculation formula for solving triangles to calculate the angle from the side lengths. Taking the calculation of angle A1 in Figure 6 as an example for illustration, the calculation formula is as follows:
[0128]
[0129] In the formula, a is the length of the opposite side line segment of the vertex where the angle A1 to be calculated is located, b and c are the lengths of the adjacent side line segments of the vertex where the angle A1 to be calculated is located respectively, and A1 is the angle of the angle to be calculated.
[0130] According to the above method, the rotation reference angles corresponding to all digital characters can be calculated. For the digital characters on the right side of the pointer center point, according to the characteristic that each scale number of the pointer-type instrument increases continuously with the rotation direction of the pointer, an angle greater than 180 degrees should be obtained. After directly calculating the angle A1 by the above formula, (360 - A1) is the real rotation reference angle.
[0131] Construct key-value pairs <digital character text content, rotation reference angle> to obtain a key-value pair sequence. Use the digital character text content as the sorting keyword to sort the key-value pair sequence in ascending order. Check whether the rotation reference angle keywords of each key-value pair after sorting also conform to ascending order. If there are rotation reference angles in the key-value pair sequence that do not conform to ascending order, the digital characters corresponding to these key-value pairs are abnormal digital characters, and they are determined as non-scale digital characters and removed.
[0132] The digital characters retained after screening through the above steps are the final scale number recognition results, hereinafter simply referred to as scale numbers.
[0133] The matching part of the pointer end point location interval matches the scale number coordinates based on the distance and angle relationship between the pointer end point and the scale number to determine the scale interval where the pointer end point is located. The matching principle is: Use the straight-line distance from the scale number to the pointer end point as the matching index. Under the condition that the end point is located between the upper and lower boundary numbers of the scale interval, take the scale number with the shortest straight-line distance to the pointer end point as the matching result. The specific matching process includes:
[0134] 1) Calculate the Euclidean distance from the end point of the pointer to each scale number, and take the two scale numbers with the smallest distance as the upper and lower bound scale numbers of the pending scale interval;
[0135] 2) For the convenience of explanation, define the concept of the angle between two points. In this patent, the angle between two points refers to the angle of the acute angle formed with the center point of the pointer rotation as the vertex and the line segments from the two points to the center point of the pointer rotation as the two sides. Calculate the angles between the end point of the pointer and the two pending scale numbers respectively, as well as the angle between the two scale numbers. The angle calculation method is exactly the same as the method of calculating the angle value based on three points described in the pointer scale interval matching module part;
[0136] 3) If the sum of the angles from the end point of the pointer to the two scale numbers is equal to the angle between the two scale numbers, it means that the pointer is indeed between the two scale numbers and the matching is successful. If not, introduce the scale number closest to the end point of the pointer among the remaining digital characters to form pending scale intervals with the above two digital scales respectively, and repeat the above matching process until the matching is successful. Among the successfully matched digital characters, the character with the larger digital character text content is called the upper bound number of the scale interval, and the character with the smaller digital character text content is called the lower bound number of the scale interval.
[0137] After matching, the effective information of the instrument reading is obtained: the end point of the pointer, the center point of the pointer rotation, and the upper and lower bound numbers of the scale interval where the pointer is located.
[0138] Such as Figure 7 Shown is a schematic diagram of an example of instrument reading calculation. In the figure, point F represents the lower bound number of the scale interval of the instrument, point G represents the upper bound number of the scale interval of the instrument, point O' represents the center point of the pointer rotation, and point E represents the end point of the pointer.
[0139] The said instrument reading module is improved from the angle method, that is, by calculating the rotation angle of the pointer and comparing it with the total angle of the range to obtain the specific reading. The formula is expressed as follows:
[0140]
[0141] In the formula, Reading represents the finally obtained instrument reading; max_scale and min_scale respectively represent the digital character text contents of the upper bound number and the lower bound number of the scale interval, corresponding to the digital character text content "20" of point F and the digital character text content "40" of point G recognized respectively; ang_end represents the angle between the center point of the pointer rotation and the lower bound number of the scale interval, corresponding to ∠FO'E in the schematic diagram; ang_interval represents the angle between the upper bound number and the lower bound number of the scale interval, corresponding to angle ∠FO'G in the schematic diagram;
[0142] Finally, the calculated readings of the instrument in this example are as follows:
[0143]
[0144] The above embodiments are preferred embodiments of the present invention. However, the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An automatic reading method for circular pointer-type instruments based on deep learning, characterized in that, It includes the following steps: 1) Obtain the images of the inspection site through a visible light camera, which are called scene images; 2) Use a dial area extraction module developed based on deep learning instance segmentation technology to detect the instrument dial area in the scene image, segment and extract the instrument dial image from the scene image, remove the background part, and the obtained part that only contains the instrument dial is called the dial ROI image; 3) By comprehensively considering the size and blur degree of the dial ROI image, obtain an index for measuring the readability of the instrument, which is called the instrument readability index. According to the calculation result of this index, match the corresponding image enhancement module to perform image enhancement on the dial ROI image; 4) Use a dial tilt correction module developed based on image transformation technology to solve the problem of dial tilt caused by the camera not being directly facing the plane where the instrument is located during the shooting process. The dial image after image enhancement and dial correction is called the high-quality dial ROI image; 5) Use deep learning technology to extract instrument dial information. Respectively use the dial character information extraction module developed based on deep learning OCR technology to obtain the character information of the high-quality dial ROI image and the dial pointer information extraction module developed based on deep learning key point detection technology to obtain the pointer information of the high-quality dial ROI image; 6) Use a pointer scale interval matching module to obtain the effective information of the instrument reading from the extraction results of the above instrument dial information. The effective information of the instrument reading respectively refers to: the end point of the pointer, the rotation center point of the pointer, and the upper and lower bound scale numbers of the scale interval where the pointer is located; 7) Use an instrument reading module to calculate the instrument reading based on the above effective information of the instrument reading; the instrument reading module is improved from the angle method, and the formula is expressed as follows: In the formula, Reading represents the finally obtained instrument reading; max_scale and min_scale respectively represent the digital character text contents of the upper bound number and the lower bound number of the scale interval, which can be directly obtained after matching and determining the upper and lower bound scale numbers; ang_end represents the angle between the rotation center point of the pointer and the lower bound number of the scale interval; ang_interval represents the angle between the upper bound number and the lower bound number of the scale interval; for the angle calculation of any three non-collinear point coordinates, as follows: Use the Pythagorean theorem formula to calculate the line segment length between two points: Where a' represents the length of the line segment, (x b , y b ) and (x c , y c ) are the endpoint coordinates of both ends of the line segment respectively. Based on the above formula, the lengths of the three sides of a triangle formed by any three non-collinear points can be calculated; Use the arbitrary angle calculation formula of solving triangles to calculate the angle from the side lengths: In the formula, a is the length of the opposite side line segment of the vertex where the angle to be calculated is located, b and c are respectively the lengths of the adjacent side line segments of the vertex where the angle to be calculated is located, and A1 is the angle to be calculated; The result calculated by the above formula is the final result of the automatic reading recognition of the circular pointer type instrument.
2. The automatic reading method for circular pointer-type instruments based on deep learning according to claim 1, wherein: In step 2), the dial area extraction module is a trained instance segmentation deep learning model. Through this model, the rectangular bounding box coordinates and image mask of the instrument dial in the scene image can be obtained. Then, the sub-image within the range of the above-mentioned rectangular bounding box is segmented from the scene image, and the part outside the dial area in the above sub-image is removed using the image mask, obtaining the part that only contains the instrument dial, which is called the dial ROI image.
3. The automatic reading method of a circular pointer-type instrument based on deep learning according to claim 2, wherein: In step 3), the instrument readability index is denoted as F, and the calculation formula of this index is as follows: In the formula, var represents the image variance, which describes the blurriness of the image; H and W respectively represent the height and width of the image, which describe the size of the image; σ and β are two weight coefficients used to adjust the influence weights of blurriness and size on the calculation result of the instrument readability index. Among them, the image variance var represents the calculation result of summing the squares of the differences between the gray values of all pixels and the average gray value of the image, and then dividing by the total number of pixels. The specific calculation process includes the following steps: 3.1) Convert the color dial ROI image into a dial ROI grayscale image. 3.2) Perform convolution operation on the dial ROI grayscale image using the following Laplacian operator template: 3.3) Calculate the variance σ of the grayscale image of the dial ROI after convolution using the following formula 2 :[[]]END]] In the formula, x represents the pixel value of a point on the image, μ represents the gray mean of the image, the summation range of the numerator includes all pixels in the image, and W*H represents the total number of pixels in the image.
4. A method for automatically reading a circular pointer-type instrument based on deep learning according to claim 3, characterized in that: In step 3), the image enhancement module includes a super-resolution reconstruction deep learning model and a bilateral filtering module. According to different calculation results of the instrument readability index, different image enhancement methods are selected to enhance the dial ROI image. According to the instrument readability index F of the input dial ROI image, if F is greater than a preset threshold, it is determined that the dial ROI image is a "low readability image", and the super-resolution reconstruction deep learning model is used to perform image super-resolution reconstruction processing on the input image; if F is less than the preset threshold, it is determined that the dial ROI image is a "non-low readability image", and the bilateral filtering module is used to perform image enhancement processing on the image. The calculation formula of the bilateral filtering module is: In the formula, f(x, y) is the response at the position of the pixel point (x, y) after filtering, g(x, y) is the value of each pixel point within the neighborhood of the pixel point (x, y), and W(x, y) is the combined weight coefficient of each pixel point; the function of this bilateral filtering module is to perform smoothing filtering on the rest of the part while keeping the edges clear to remove image noise and improve image quality.
5. A method for automatically reading a circular pointer-type instrument based on deep learning according to claim 4, characterized in that: In step 4), the dial tilt correction module specifically performs the following operations: 4.1) Obtain the instrument image mask obtained by the dial area extraction module, which is the binary mask mask image of the instrument contour. 4.2) Perform ellipse fitting on the above mask image to obtain the contour fitting ellipse parameters of the dial ROI image, including the ellipse center point, the length of the major axis of the ellipse, the length of the minor axis of the ellipse, and the angle between the major axis of the ellipse and the vertical direction. 4.3) Obtain the coordinates of the endpoints of the major and minor axes of the ellipse obtained after fitting, and calculate the coordinates of the points on the line where the minor axis is located and the distance to the center point of the ellipse is the length of the major axis. These coordinates are called the corrected expected coordinates of the minor axis of the ellipse; 4.4) Use the coordinates of the endpoints of the major and minor axes of the ellipse, the coordinates of the endpoints of the major axis of the ellipse, and the corrected expected coordinates of the minor axis of the ellipse to form 4 sets of feature point pairs. Use these feature point pairs to calculate the projective transformation matrix. The calculation of the projective transformation matrix is as follows: p2 = H′ * p1 In the formula, p1 represents a point in the original image before transformation, p2 represents the corresponding feature point of p1, and H′ is the projective transformation matrix, which represents the process of mapping the point p1 in the original image to the point p2 in the transformed image. Expanding the above formula gives: where (x1, y1) represents the coordinates of point p1, (x2, y2) represents the coordinates of point p2, and H 11 ~H 33 are all parameters of matrix H′, and a system of equations can be obtained from the above matrix form; among them, from H 11 ~H 33 these 9 parameters can be used to construct a system of equations with the above 4 groups of feature point pairs to solve for the unique projective transformation matrix H′; 4.5) The projective transformation matrix H′ obtained by using the formula in step 4.4) can perform global processing on the dial ROI image, calculate the pixel coordinates of each dial ROI image after correction before correction, so as to realize the correction of the elliptical dial ROI image into a circular shape.
6. The automatic reading method for circular pointer type meters based on deep learning according to claim 5, characterized in that: In step 5), the dial character information extraction module needs to extract the dial character information including the character text content and its corresponding position coordinates of each character image. The acquisition process is specifically including the following steps: 5.1.1) Input the high-quality dial ROI image that has undergone image enhancement and dial correction into a trained deep learning model for scene text detection. The model infers and outputs all character region information of the high-quality dial ROI image. The character region information is the rectangular bounding box of the character, which is described by the four vertex coordinates of the rectangular bounding box; 5.1.2) Sequentially segment all characters from the high-quality dial ROI image through the rectangular bounding box coordinates of the character images to obtain a sequence of character images of the high-quality dial ROI image; 5.1.3) Use the vertex coordinates of the rectangular bounding box of each character image to calculate the character position coordinates. The character position coordinates are defined as the center coordinates of the vertices of the character rectangular bounding box, and their horizontal and vertical coordinates are the averages of the four vertex coordinates of the rectangular bounding box; 5.1.4) Sequentially input each character image in the sequence of character images into a trained deep learning model for text recognition of character images to obtain the text content recognition results of each character image in the sequence of character images. Save the text content and the corresponding position coordinates in pairs, and the dial character information is obtained.
7. A method for automatically reading a circular pointer-type instrument based on deep learning according to claim 6, characterized in that: In step 5), the dial pointer information extraction module specifically performs the following operations: 5.2.1) Detect the pointers inside the dial through a deep learning model for object detection to obtain the rectangular bounding box parameters of the pointers, including the position coordinates of the center point of the pointer and the width and height of the rectangular bounding box; 5.2.2) Based on the rectangular bounding box parameters of the pointer, obtain the sub-image inside the range of the pointer rectangular bounding box from the scene image. The obtained image is called the pointer ROI image; 5.2.3) Input the pointer ROI image into a trained deep learning model for key point detection. Through the model, directly infer the position coordinates of the pointer rotation center point and the pointer end point. The above two points and their positions are the extracted pointer information.
8. A method for automatically reading a circular pointer-type instrument based on deep learning according to claim 7, characterized in that: In step 6), the pointer scale interval matching module is divided into two stages: scale number screening and matching the scale interval where the pointer is located: Stage 1, scale number screening: 6.1.1) Define the scale numbers as the numeric characters representing the instrument scales in the dial ROI image. First, filter out all the characters with pure numeric text content from the character text content obtained by the dial information extraction module as the preliminary screening results of the scale numbers, hereinafter referred to as numeric characters; 6.1.2) Calculate the Euclidean distances from the above-screened numeric characters to the pointer rotation center point to obtain a distance sequence; 6.1.3) Use the k-means clustering algorithm to cluster the above distance sequence. Set the number of clustering categories k to 3. Since the scale numbers of the pointer-type instrument are approximately distributed on the same circle in terms of spatial distribution, the distances to the rotation center point are approximately the same and are clustered into the same category. The numeric characters within this category are used as the effective screening results of the numeric characters, and the remaining numeric characters are determined as invalid screening results and removed from the numeric characters; 6.1.4) Calculate the rotation reference angles of the remaining numeric characters. The rotation reference angle is defined as the angle between the line connecting the numeric character and the pointer rotation center point and the ray vertically downward from the pointer rotation center point with the line passing through the pointer rotation center point and pointing downward as the reference line; 6.1.5) Use the numeric character text content and the corresponding rotation reference angle to form a <numeric character text content, rotation reference angle> key-value pair. The key-value pairs of all the remaining numeric characters form a key-value pair sequence; 6.1.6) Use the numeric character text content as the sorting keyword to sort the above key-value pair sequence in ascending order. Check whether the rotation reference angles of the key-value pairs in the sorted key-value pair sequence also conform to the ascending order rule. If there are key-value pairs whose rotation reference angles do not conform to the ascending order rule, the numeric characters corresponding to these key-value pairs are determined as non-scale numeric characters and removed from the numeric characters; 6.1.7) The characters retained after the above screening are the final scale numbers; Stage 2, matching the scale interval where the pointer is located: 6.2.1) Calculate the Euclidean distances from the pointer end point to each scale number, and use the two scale numbers with the closest distances to the pointer end point as the upper and lower bound scale numbers of the scale interval; 6.2.2) Define the angle between two points as the acute angle formed by taking the pointer rotation center point as the vertex and the lines connecting the two points to the pointer rotation center point as the two sides, where the two points are the upper and lower bound scale numbers; According to the above definition, the angles between the pointer end point and the upper and lower bound scale numbers of the to-be-determined scale interval and the angle between the upper and lower bound scale numbers of the to-be-determined scale interval can be calculated; 6.2.3) If the sum of the angles between the end point and the scale numbers at the upper and lower bounds of the to-be-determined scale interval is approximately equal to the angle between the scale numbers at the upper and lower bounds of the to-be-determined scale interval, the matching is successful; if not, introduce the digital character closest to the end point of the pointer among the remaining digital characters to form to-be-determined scale intervals with the above two scale numbers respectively, and repeat the above matching process until the matching is successful. Among the successfully matched digital characters, the character with a larger digital character text content is called the upper bound number of the scale interval, and the character with a smaller digital character text content is called the lower bound number of the scale interval.
Citation Information
Patent Citations
Automatic identification reading method and system for pointer instrument
CN112818988A
Intelligent inspection pointer type instrument identification and reading method based on deep learning
CN114549981A