A Pointer Instrument Recognition Method Based on Text Region Reading

The method uses a convolutional neural network to detect gauge text, correct image angles, and convert to polar coordinates for flexible and precise reading of gauges with minimal training, addressing inefficiencies in traditional gauge reading methods.

CN114399677BActive Publication Date: 2025-07-15黄丽莉
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111599434.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-07-15
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

Traditional methods are inefficient when reading instruments, cannot flexibly adapt to different types of instruments, and require a large amount of training data for retraining, making it difficult to apply in different types of meters.

Method used

The algorithm based on convolutional neural network is used to detect the tick value text, use the text border center coordinates to perform image correction, and perform polar coordinate transformation. The pointer and tick line are extracted in combination with the quadratic region search method, and read the readings using the distance method.

Benefits of technology

High-precision and robust text recognition are achieved, reducing the impact of shooting angle on recognition, adapting to different types of instruments, reducing training data requirements, and improving reading accuracy and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399677B_ABST
    Figure CN114399677B_ABST
Patent Text Reader

Abstract

The present invention discloses a pointer instrument recognition method based on text area reading. The recognition method involves an algorithm, and the algorithm includes the following processes: First, the algorithm applies deep learning to the detection of scale value text. After obtaining an accurate text border, image correction is performed using the center coordinates within the text border to eliminate the influence of the shooting angle on reading and recognition. Then, the meter center is determined according to the center point in the text bounding box, and polar coordinate transformation is performed to convert the circular scale line into a horizontal scale line. On this basis; finally, the reading between the two main scale lines closest to the pointer is obtained by the distance method. The present invention applies deep learning to the detection of instrument scale value text, achieving high precision and robustness in text coordinate positioning and high precision in text recognition. In addition, compared with the distance method of reading from the zero scale to the maximum scale, the recognition result of the scale value read by the distance method allows for a smaller error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a pointer instrument recognition method based on text area reading. Background Art

[0002] For the traditional method of reading instruments by the angle method, it is necessary to manually input the angles of the zero scale line and the full scale line in the image, and the efficiency is low. Using the slope of the instrument edge line to correct the instrument image is only applicable to the instrument panel with a straight edge. And when directly reading the electric meter by using an end-to-end neural network model, when a new type of instrument appears, a large amount of training data needs to be prepared again, and the network needs to be retrained. The algorithm cannot be flexibly used on different instruments. And when using Mask-RCNN to obtain the pointer area, for a new instrument with different pointer shapes, a large amount of data sets need to be prepared and the network needs to be retrained. This algorithm cannot be flexibly applied to different types of meters. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides the following technical solutions:

[0004] The present invention provides a pointer instrument recognition method based on text area reading.

[0005] The recognition method involves an algorithm that uses a convolutional neural network to detect the scale value text. The algorithm includes the following processes:

[0006] First, the algorithm applies deep learning to the detection of scale value text. After obtaining the accurate text border, it uses the central coordinates within the text border to correct the image, thereby eliminating the influence of the shooting angle on reading and recognition.

[0007] Then, it determines the meter center according to the center point in the text bounding box, performs polar coordinate transformation to convert the circular scale line into a horizontal scale line. On this basis, a secondary region search method is proposed to extract the pointer and scale line images of the instrument.

[0008] Finally, the reading between the two main scale lines closest to the pointer is obtained by the distance method.

[0009] Preferably, it specifically includes the following steps:

[0010] Step 1: Collect instrument images, and use the FOTS shared convolutional neural network to detect and recognize the scale value text in the instrument. The FOTS shared convolutional neural network is an end-to-end text recognition model. It uses the shared convolutional network to extract the shared features of the collected instrument images, and uses these features to determine the position of the text area in the detection part; according to the position of the text area detected by the detection part, extract this part of the features from the shared features for predicting and recognizing the text in this area.

[0011] Use ResNet50 for image encoding, and then decode to obtain shared features through repeated upsampling, concatenation operations, and two-layer convolution. For the points on the feature map, in the detection part, first predict whether these points belong to the text region, then predict the distances from these points to the four boundaries of the text region and the rotation angle of the text box, then set a threshold to filter the points, and perform non-maximum suppression on the generated prediction boxes. Finally, obtain multiple regions. The RoIRotate module uses bilinear interpolation sampling to convert the text feature map with an indefinite length and a determined angle into a feature map with a determined length and no angle;

[0012] Step 2: Correct the instrument image; to reduce subsequent reading errors, here, through the method of SIFT feature matching, extract the feature points of the instrument image and the template image, screen the feature points through the RANSAC algorithm and establish a homography matrix, obtain the mapping of the vertices of the text box detected in Step 1, and use perspective transformation to correct the instrument image;

[0013] Step 3: After locating the scale value text and image correction, determine the center of polar coordinate transformation and expand the polar coordinates; the scale distribution of pointer-type instruments with high precision is often relatively dense, making it difficult to separate single scales in the curve area, and the same is true for the application of the angle method in reading. Therefore, it is necessary to perform polar coordinate transformation on the instrument image to convert the curve scale into a linear scale whose relative position is easy to calculate. The polar coordinate transformation includes a center extraction method, which is to extract the center using the text bounding box. The scale values of the instrument are distributed on an arc, and the center of the arc is the rotation center of the instrument pointer. Therefore, by fitting the arc with the coordinates of the scale value text as data points, the center of polar coordinate transformation can be determined;

[0014] Step 4: Extract the pointer and scale lines; after calculating the polar coordinate radius and polar coordinate angle of each pixel in the polar coordinate system, expand them in the rectangular coordinate system with the polar coordinate radius ρ and polar coordinate angle θ as the abscissa and ordinate. The expansion formula is as follows: where x0 and y0 are the abscissa and ordinate in the original coordinate system, C x Cy is the pole in the polar coordinate system,

[0015]

[0016]

[0017] Step 5: Perform Otsu threshold segmentation on the instrument image and convert it into a binary image. Cut the ROI area according to the obtained rectangular coordinates, and then find the accumulated number of pixels in the projected horizontal position of each column of white pixel images. The pointer after projection is the position with the least number of pixels, that is, the horizontal position where the pointer is located. Although the shape of the pointer varies with the type of instrument, the number of projected pixels must be the minimum number of pixels at the horizontal position of the pointer, which is always consistent. Therefore, the horizontal coordinate of the pointer obtained by projection has high accuracy and robustness.

[0018] Step 6: Extraction of scale lines. Compared with the pointer, the features of scale lines are not so obvious and are more easily affected by other objects in the dial. Therefore, the search range is further narrowed on the basis of the primary search area to extract the scale lines. According to the position of the scale value text, the secondary search area can be obtained by the following steps: the area containing the scale value text border can be obtained according to the vertex and then the same area is formed above it. Another vertical projection is performed in the secondary search area. The white pixels of each row in the image are projected onto the x-axis to find the horizontal position with the least cumulative number of pixels. In the pixel coordinate system of the ROI area, the horizontal coordinate X of the pointer is calculated. pointer , and the horizontal coordinate X of the scale line corresponding to the scale value text on both sides of the pointer l-scale and X r-scale , as shown below

[0019]

[0020] Where V is the final reading; V r is the scale value corresponding to the scale line on the right side of the pointer, V l It is the scale value corresponding to the scale line on the left side of the pointer.

[0021] Preferably, the polar coordinate transformation is to transform an image from a Cartesian coordinate system into a polar coordinate system centered at a certain point in the image.

[0022] Preferably, in the step 1, in the recognition part, the input feature map is first further encoded using a convolutional neural network that shrinks only in height, and then the features are decoded using a bidirectional LSTM to generate a final predicted string.

[0023] Preferably, in step three, after obtaining the precise coordinates of the scale value text box, the center coordinates are obtained by least square fitting.

[0024] Preferably, the convolutional neural network is used to detect the scale value text. Since the scale value text in the instrument is the same, the network training can be completed with a small sample data. When a new type of instrument appears, the algorithm can also show good robustness.

[0025] Advantages of the present invention

[0026] (1) The present invention applies deep learning to the detection of instrument scale value text, achieving high precision and robustness in text coordinate positioning and high precision in text recognition. In addition, compared with the distance method of reading from zero scale to the maximum scale, the recognition result of the scale value read by the distance method allows a smaller error.

[0027] (2) The present invention proposes a new method for positioning the center of the pointer. This method locates the center of the circle according to the position of the scale value text. The scale value text image provides more features than the scale line image, so it can adapt to more complex environments when used to fit the scale curve.

[0028] (3) The present invention proposes a secondary region search method based on the position of the scale value text to extract the pointer and the scale line. This method effectively solves the problem of pointer shadow and also eliminates the influence of other objects on the scale dial on the extraction of the pointer and the scale line. Description of the drawings

[0029] Figure 1 It is a flowchart of an algorithm step provided by the present invention.

[0030] Figure 2 It is a schematic diagram of the FOTS shared convolutional neural network structure of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] Embodiment

[0033] As Figure 1 shown, the present invention provides a pointer instrument recognition method based on text region reading.

[0034] The recognition method involves an algorithm that uses a convolutional neural network to detect scale value text. Since the scale value text in the instrument is the same, the network can be trained with small sample data. When a new type of instrument appears, the algorithm can also show good robustness. The algorithm includes the following processes:

[0035] First, the algorithm applies deep learning to the detection of scale value text. After obtaining the accurate text bounding box, it uses the central coordinates within the text bounding box for image correction to eliminate the influence of the shooting angle on reading and recognition.

[0036] Then, it determines the meter center based on the center point in the text bounding box, performs polar coordinate transformation to convert the circular scale lines into horizontal scale lines. On this basis, a quadratic region search method is proposed to extract the pointer and scale line images of the meter.

[0037] Finally, the reading between the two main scale lines closest to the pointer is obtained by the distance method.

[0038] Specifically, it includes the following steps:

[0039] Step 1: Collect the meter image, and use the FOTS shared convolutional neural network to detect and recognize the scale value text in the meter. The FOTS shared convolutional neural network is an end-to-end text recognition model. It uses the shared convolutional network to extract the shared features of the collected meter image, and uses these features to determine the position of the text area in the detection part. According to the position of the text area detected in the detection part, extract this part of the features from the shared features for predicting and recognizing the text in this area.

[0040] Use ResNet50 for image encoding, and then decode to obtain the shared features through repeated upsampling, concatenation operations, and two layers of convolution. For the points on the feature map, in the detection part, first predict whether these points belong to the text area, then predict the distances from these points to the four boundaries of the text area and the rotation angle of the text box, then set a threshold to filter the points, perform non-maximum suppression on the generated prediction boxes, and finally obtain multiple regions. The RoIRotate module uses bilinear interpolation sampling to convert the text feature map with variable length and determined angle into a feature map with determined length and no angle. In the recognition part, first use a convolutional neural network that only shrinks in height to further encode the input feature map, and then use a bidirectional LSTM to decode the features to generate the final predicted string.

[0041] Step 2: Correct the meter image; to reduce subsequent reading errors, here, the feature points of the meter image and the template image are extracted by the SIFT feature matching method, the feature points are screened by the RANSAC algorithm and a homography matrix is established, and the mapping of the vertices of the text box detected in Step 1 is obtained, and the meter image is corrected using perspective transformation.

[0042] Step 3: After positioning the scale value text and correcting the image, determine the polar coordinate transformation center and expand the polar coordinates; the scale distribution of pointer-type instruments with higher precision is often dense, which makes it difficult to separate a single scale in the curve area, and the same is true for the application of the angle method in reading. Therefore, it is necessary to perform polar coordinate transformation on the instrument image to convert the curve scale into a linear scale whose relative position is easy to calculate. The polar coordinate transformation is to convert an image from a Cartesian coordinate system to a polar coordinate system centered on a certain point in the image. The polar coordinate transformation includes a center extraction method, which uses a text bounding box to extract the center. The scale values of the instrument are distributed on an arc, and the center of the arc is the rotation center of the instrument pointer. Therefore, the coordinates of the scale value text are used as data points to fit the arc, so that the center of the polar coordinate transformation can be determined. After obtaining the precise coordinates of the scale value text box, the least squares method is used to fit the center coordinates;

[0043] Step 4: Extract the pointer and scale line; after calculating the polar coordinate radius and polar coordinate angle of each pixel in the polar coordinate system, use the polar coordinate radius ρ and polar coordinate angle θ as the horizontal coordinate and vertical coordinate, and expand it in the rectangular coordinate system. The expansion formula is as follows: where x0, y0 are the horizontal coordinate and vertical coordinate in the original coordinate system, C x Cy is the pole in the polar coordinate system,

[0044]

[0045]

[0046] Step 5: Perform Otsu threshold segmentation on the instrument image and convert it into a binary image. Cut the ROI area according to the obtained rectangular coordinates, and then find the accumulated number of pixels in the projected horizontal position of each column of white pixel images. The pointer after projection is the position with the least number of pixels, that is, the horizontal position where the pointer is located. Although the shape of the pointer varies with the type of instrument, the number of projected pixels must be the minimum number of pixels at the horizontal position of the pointer, which is always consistent. Therefore, the horizontal coordinate of the pointer obtained by projection has high accuracy and robustness.

[0047] Step 6: Extraction of scale lines. Compared with the pointer, the features of scale lines are not so obvious and are more easily affected by other objects in the dial. Therefore, the search range is further narrowed on the basis of the primary search area to extract the scale lines. According to the position of the scale value text, the secondary search area can be obtained by the following steps: the area containing the scale value text border can be obtained according to the vertex and then the same area is formed above it. Another vertical projection is performed in the secondary search area. The white pixels of each row in the image are projected onto the x-axis to find the horizontal position with the least cumulative number of pixels. In the pixel coordinate system of the ROI area, the horizontal coordinate X of the pointer is calculated. pointer, and the horizontal coordinate X of the scale line corresponding to the scale value text on both sides of the pointer l-scale and X r-scale , as follows

[0048]

[0049] wherein, V is the final reading; V r is the scale value corresponding to the scale line on the right side of the pointer, and V l is the scale value corresponding to the scale line on the left side of the pointer.

[0050] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A pointer instrument recognition method based on text area reading, characterized in that the recognition method involves an algorithm that uses a convolutional neural network to detect scale value text, and the algorithm includes the following processes: First, the algorithm applies deep learning to the detection of scale value text. After obtaining an accurate text bounding box, it uses the central coordinates within the text bounding box for image correction to eliminate the influence of the shooting angle on reading and recognition. Then, it determines the meter center based on the center point in the text bounding box, performs polar coordinate transformation to convert the circular scale line into a horizontal scale line. On this basis, a secondary region search method is proposed to extract the pointer and scale line images of the instrument. Finally, the reading between the two main scale lines closest to the pointer is obtained by the distance method. The secondary region search method specifically includes the following steps: Step 1: Collect the instrument image, and use the FOTS shared convolutional neural network to detect and recognize the scale value text in the instrument. The FOTS shared convolutional neural network is an end-to-end text recognition model. It uses the shared convolutional network to extract the shared features of the collected instrument image, and uses these features to determine the position of the text region in the detection part. According to the position of the text region detected in the detection part, this part of the features is extracted from the shared features for predicting and recognizing the text in this region. Use ResNet50 for image encoding, and then decode through repeated upsampling, connection operations, and two layers of convolution to obtain shared features. For the points on the feature map, in the detection part, first predict whether these points belong to the text region, then predict the distances from these points to the four boundaries of the text region and the rotation angle of the text box, then set a threshold to filter the points, perform non-maximum suppression on the generated prediction boxes, and finally obtain multiple regions. The RoIRotate module uses bilinear interpolation sampling to convert the text feature map with an uncertain length and a determined angle into a feature map with a determined length and no angle. Step 2: Correct the instrument image; to reduce subsequent reading errors, here, the SIFT feature matching method is used to extract the feature points of the instrument image and the template image, screen the feature points through the RANSAC algorithm and establish a homography matrix, obtain the mapping of the vertices of the text box detected in Step 1, and correct the instrument image using perspective transformation. Step 3: After locating the scale value text and correcting the image, determine the polar coordinate transformation center and expand the polar coordinates; convert the curve scale into a linear scale whose relative position is easy to calculate. The polar coordinate transformation includes a center extraction method, and the center extraction method is to extract the center using the text bounding box. The scale values of the instrument are distributed on an arc, and the center of the arc is the rotation center of the instrument pointer. Therefore, by fitting the arc with the coordinates of the scale value text as data points, the center of the polar coordinate transformation can be determined. Step 4: Extract the pointer and scale lines; after calculating the polar radius and polar angle of each pixel in the polar coordinate system, using the polar radius and the polar angle as the abscissa and ordinate, expand them in the rectangular coordinate system. The expansion formula is as follows: where , are the abscissa and ordinate in the original coordinate system, is the pole in the polar coordinate system. Step 5: Perform Otsu threshold segmentation on the instrument image and convert it into a binary image. Cut the ROI area according to the obtained rectangular coordinates, and then find the accumulated number of pixels at the projected horizontal position of each column of white pixel images. The pointer after projection is the position with the least number of pixels, that is, the horizontal position of the pointer; Step 6: Extraction of scale lines. Compared with the pointer, the features of the scale lines are less obvious and are more easily affected by other objects inside the dial. Therefore, on the basis of the main search area, the search range is further narrowed to extract the scale lines. According to the position of the scale value text, the secondary search area is obtained through the following steps: The area containing the border of the scale value text is obtained according to the vertices and then an area identical to it is formed above it. Another vertical projection is performed in the secondary search area, where the white pixels of each row in the image are projected onto the x-axis to find the horizontal position with the least cumulative number of pixels. In the pixel coordinate system of the ROI area, the horizontal coordinate of this pointer is calculated , as well as the horizontal coordinates of the scale lines corresponding to the scale value texts on both sides of this pointer and , as follows Among them, V is the final reading; is the scale value corresponding to the scale line on the right side of the pointer, and is the scale value corresponding to the scale line on the left side of the pointer.

2. The method for identifying a pointer instrument based on text area reading according to claim 1, characterized in that: The polar coordinate transformation is to transform an image from a Cartesian coordinate system to a polar coordinate system centered at a certain point in the image.

3. The method for identifying a pointer instrument based on text area reading according to claim 1, characterized in that: In the step 1, in the recognition part, the input feature map is first further encoded using a convolutional neural network that shrinks only in height, and then the features are decoded using a bidirectional LSTM to generate the final predicted string.

4. The method for identifying a pointer instrument based on text area reading according to claim 1, characterized in that: In step three, after obtaining the precise coordinates of the scale value text box, the center coordinates are obtained by using the least squares fitting method.

Citation Information

Patent Citations

  • Automatic identification reading method and system for pointer instrument

    CN112818988A

  • Meter reading identification method and device, readable storage medium and computer equipment

    CN113673486A