Character and Line with Width Recognition Method for Scanned Electrical CAD Drawings

Through the combination of OCR model and traditional machine learning, the problem of low character recognition accuracy and unrecognized line width of electrical CAD drawings is solved, efficient character and line recognition is achieved, and electronic management of electrical CAD drawings is improved.

CN115100657BActive Publication Date: 2025-07-22BAOSHAN POWER SUPPLY BUREAU OF YUNNAN POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210164628.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-07-22
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

In the prior art, the character recognition accuracy of the electrical CAD drawing scan drawing is low and the line width cannot be recognized, which affects the understanding and application of the drawing.

Method used

The OCR model is used to combine conditional screening and traditional machine learning character recognition algorithms, and combine edge detection and clustering algorithms to identify characters and obtain line widths.

Benefits of technology

The character recognition accuracy rate is improved to 93%, and the line width is accurately obtained, which reduces the interference of character recognition and improves the electronic management efficiency of electrical CAD drawings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100657B_ABST
    Figure CN115100657B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying characters and lines with width in scanned electrical CAD drawings, belonging to the technical field of electrical CAD drawing recognition. After cutting the image to be recognized into images of equal size, the images are sequentially input into the OCR model for character detection; cropping is performed according to the coordinate results of character detection; the cropped character images are sent into the OCR model for recognition and screening, and the coordinate transformation of the recognized character regions is carried out to obtain the true corresponding coordinates of the character regions on the entire original image; then the characters and components are covered; line detection is performed on the covered image and the detected lines are classified; the vertical lines and horizontal lines are respectively connected and merged; clustering is respectively performed on the vertical lines, horizontal lines and oblique lines to obtain the coordinates of the lines and the corresponding line widths. The present invention solves the problems of low accuracy of character recognition in scanned electrical CAD drawings and inability to recognize line widths, and is easy to promote and apply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electrical CAD drawing recognition, and particularly relates to a method for recognizing characters and lines with width in scanned electrical CAD drawings. Background Art

[0002] In the process of converting a scanned paper electrical CAD drawing into an electronic CAD drawing, character recognition and line recognition are the two most important parts. However, the current effect of character recognition for scanned electrical CAD drawings is not very ideal. The character recognition is inaccurate and there are problems of recognition errors. For the line recognition of electrical CAD drawings, the line width in the drawing has certain significance and has a crucial impact on the understanding and application of the electrical CAD drawing. However, the existing methods for line recognition of scanned electrical CAD drawings cannot recognize the line width.

[0003] Therefore, how to improve the accuracy of character recognition for scanned electrical CAD drawings and process the line width when recognizing lines is particularly important. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems of low accuracy of character recognition and inability to recognize line width in the scanned electrical CAD drawings, and provide a method for recognizing characters and lines with width in scanned electrical CAD drawings. By using the character detection function of the OCR model and combining conditional screening with the character recognition algorithm of traditional machine learning, the correct rate of character recognition is improved; the edge detection method is used and combined with the clustering algorithm to obtain the line width of the line on the basis of line recognition.

[0005] To achieve the above purpose, the technical scheme adopted by the present invention is as follows:

[0006] A method for recognizing characters and lines with width in scanned electrical CAD drawings, comprising the following steps:

[0007] Cut the image to be recognized into images of equal size;

[0008] Sequentially input the cut images into the OCR model for character detection;

[0009] Crop according to the coordinates of the character detection result, and cut out each character image;

[0010] Send the cropped character images into the OCR model for recognition, and screen according to the recognition correct rate to obtain the recognition result; the recognition result includes the character area and the character content;

[0011] Perform coordinate conversion on the character area to obtain the true corresponding coordinates of the character area on the entire original image;

[0012] Cover the characters on the original image according to the character recognition results, and cover the components;

[0013] Perform line detection on the covered image;

[0014] Classify the detected lines, and classify the lines into vertical lines, horizontal lines, and oblique lines;

[0015] Connect and merge the vertical lines and horizontal lines respectively;

[0016] Cluster the three types of lines, namely vertical lines, horizontal lines, and oblique lines, respectively, to obtain the coordinates of the lines and the corresponding line widths.

[0017] Furthermore, preferably, cut the image to be recognized into images of equal size, and cut them into squares with a side length of 550-850 pixels.

[0018] Furthermore, preferably, send the cropped character image into the OCR model for recognition. When the accuracy rate is lower than the threshold, delete the recognition result; when the recognition accuracy rate is higher than the threshold, retain the recognition result; the threshold is not less than 80%.

[0019] Furthermore, preferably, the covered image uses fast line detection based on OpenCV to obtain line data, including the x and y values of the starting coordinates of the line and the x and y values of the ending coordinates of the line.

[0020] Furthermore, preferably, determine the smaller x value as the x value of the starting coordinate, and determine the larger x value as the x value of the ending coordinate, and modify the y value according to the corresponding relationship.

[0021] Furthermore, preferably, when classifying the detected lines, first calculate the offsets of the x and y values; subtract the x value of the starting coordinate from the x value of the ending coordinate of the line to obtain the offset of the x value, and subtract the y value of the starting coordinate from the y value of the ending coordinate of the line to obtain the offset of the y value;

[0022] If the x coordinate offset is greater than the threshold and the y coordinate offset is less than the threshold, the line is a horizontal line;

[0023] If the y coordinate offset is greater than the threshold and the x coordinate offset is less than the threshold, the line is a vertical line;

[0024] If the offsets of both the x coordinate and the y coordinate are greater than the threshold, the line is an oblique line; store the classified line data separately.

[0025] Furthermore, preferably, the threshold value ranges from 5 to 30.

[0026] Furthermore, preferably, the three types of lines, i.e., vertical lines, horizontal lines, and oblique lines, are clustered respectively to obtain the coordinates of the lines and the corresponding line widths; the specific method is:

[0027] Vertical lines are clustered according to the x-coordinate, and the line width is determined by calculating the difference between the x-coordinates of the two lines. Horizontal lines are clustered according to the y-coordinate, and the line width is determined by calculating the difference between the y-coordinates of the two lines. When clustering oblique lines, the coordinates of the center point of the line are calculated, and these coordinates are used to calculate clustering and line width.

[0028] Furthermore, preferably, when clustering, the distance between the centroids of the lines is used as the basis for clustering. When the distance between the centroids of the lines is less than a threshold, the two lines are identified as two sides of a line, that is, they are clustered into one line. The distance between the centroids is the width of the line after clustering, and the threshold is set to 10 to 20.

[0029] In the present invention, the coordinates of the lines are represented by the pixel coordinates corresponding to the lines on the picture.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] In view of the problem of low character recognition accuracy, the present invention proposes a method for character and line recognition with width in electrical CAD drawing scans. The method is based on the character selection technology of the character detection function of the OCR model, and the character selection box is processed. The selected characters are cropped and then combined with machine learning (OCR model) for character recognition. The accuracy of the character recognition algorithm is used as a screening index for character recognition to screen the recognition results. On the basis of character recognition, for the problem of line width, a method based on line edge detection and combined with threshold screening clustering is proposed to process, thereby obtaining a method for obtaining line width, and the line width of the recognized line is quickly obtained on the basis of line recognition. The method of the present invention can greatly improve the accuracy of character recognition, and can increase the accuracy to 93%, which can increase the accuracy by 3% compared with the existing recognition technology. On the basis of improving the accuracy of character recognition, the method of the present invention also improves the detection accuracy of lines and reduces the interference of characters on the detection of short lines.

[0032] By using the present invention, the necessary conditions in terms of characters and lines are provided for realizing the conversion of paper electrical CAD scanned drawings into electronic CAD drawings with low manual intervention. The utilization rate of paper electrical drawings is improved, and the foundation is laid for the management and application of electronic electrical CAD drawings. The accuracy and efficiency of recognition conversion are improved in terms of characters and lines. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a structural diagram of a method for recognizing characters and lines with widths in a scanned image of an electrical CAD drawing of the present invention;

[0034] Figure 2 It is the timing diagram of the character and line width line recognition method for the scanned electrical CAD drawing of the present invention;

[0035] Figure 3 It is the flow chart based on character recognition in the character and line width line recognition method for the scanned electrical CAD drawing of the present invention;

[0036] Figure 4 It is the flow chart of line recognition and line width calculation in the character and line width line recognition method for the scanned electrical CAD drawing of the present invention;

[0037] Figure 5 It is the character detection result annotation diagram in the application example;

[0038] Figure 6 It is the character cutting diagram in the application example;

[0039] Figure 7 It is the annotation diagram of the character recognition result on the original drawing in the application example;

[0040] Figure 8 It is the diagram after characters and components are covered in the application example;

[0041] Figure 9 It is the line clustering result table in the application example;

[0042] Figure 10 It is the annotation diagram of the line recognition result on the original drawing in the application example. Detailed implementation manner

[0043] The present invention will be further described in detail below in conjunction with the embodiments.

[0044] Those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the embodiments regarding specific technologies or conditions, they shall be carried out according to the technologies or conditions described in the literature in this field or according to the product instructions. For those materials or equipment not specified as the manufacturer, they are all conventional products that can be obtained by purchase.

[0045] The method for recognizing characters and lines with width in the scanned electrical CAD drawing includes the following steps:

[0046] Cut the picture to be recognized into pictures of equal size;

[0047] Sequentially input the cut pictures into the OCR model for character detection;

[0048] Crop according to the coordinates of the character detection results, and cut out individual character pictures;

[0049] Send the cropped character images into the OCR model for recognition, and screen them according to the recognition accuracy rate to obtain the recognition results; the recognition results include the character area and the character content;

[0050] Perform coordinate transformation on the character area to obtain the true corresponding coordinates of the character area on the entire original image;

[0051] Cover the characters on the original image according to the character recognition results, and cover the components;

[0052] Perform line detection on the covered image;

[0053] Classify the detected lines, and divide the lines into vertical lines, horizontal lines, and oblique lines;

[0054] Connect and merge the vertical lines and horizontal lines respectively;

[0055] Cluster the three types of lines, namely vertical lines, horizontal lines, and oblique lines, respectively, to obtain the coordinates of the lines and the corresponding line widths.

[0056] Preferred solution: Cut the image to be recognized into images of equal size, and cut them into squares with a side length of 550-850 pixels.

[0057] Preferred solution: Send the cropped character images into the OCR model for recognition. When the accuracy rate is lower than the threshold, delete the recognition results; when the recognition accuracy rate is higher than the threshold, retain the recognition results; the threshold is not less than 80%.

[0058] Preferred solution: Use fast line detection based on OpenCV for the covered image to obtain line data, including the x and y values of the starting coordinates of the line and the x and y values of the ending coordinates of the line.

[0059] Preferred solution: Determine the smaller x value as the x value of the starting coordinate, and determine the larger x value as the x value of the ending coordinate, and modify the y value according to the corresponding relationship.

[0060] Preferred solution: When classifying the detected lines, first calculate the offsets of the x and y values; subtract the x value of the starting coordinate from the x value of the ending coordinate of the line to obtain the x value offset, and subtract the y value of the starting coordinate from the y value of the ending coordinate of the line to obtain the y value offset;

[0061] If the x coordinate offset is greater than the threshold and the y coordinate offset is less than the threshold, the line is a horizontal line;

[0062] If the y coordinate offset is greater than the threshold and the x coordinate offset is less than the threshold, the line is a vertical line;

[0063] If the offsets of both the x - coordinate and the y - coordinate are greater than the threshold value, the line is a slant line; the classified line data are stored separately.

[0064] In a preferred embodiment, the threshold value ranges from 5 to 30.

[0065] In a preferred embodiment, clustering is performed on the three types of lines: vertical lines, horizontal lines, and slant lines, to obtain the coordinates of the lines and the corresponding line widths; the specific method is as follows:

[0066] Vertical lines are clustered according to the x - coordinate, and the difference between the x - coordinates of two lines is calculated to determine the line width. Horizontal lines are clustered according to the y - coordinate, and the difference between the y - coordinates of two lines is calculated to determine the line width. When clustering slant lines, the central point coordinates of the line are calculated, and clustering and line - width calculation are performed based on these coordinates.

[0067] In a preferred embodiment, when clustering, the distance between the centroids of the lines is used as the basis for clustering judgment. When the distance between the centroids of the lines is less than the threshold value, the two lines are considered as the two sides of one line, that is, they are clustered into one line, and the distance between the centroids is the width of the clustered line. The threshold value is set to 10 - 20.

[0068] The OCR model described in the present invention is an existing open - source model, which can be retrained by providing training data according to the recognition requirements to obtain a higher recognition accuracy, or can be directly used for recognition by the OCR model, but the recognition accuracy will be insufficient. OpenCV is an existing open - source Python library.

[0069] The picture is cropped so that the pixel sizes of the length and width of the picture are between 550 and 850, which can reduce the missed detection situation of character detection and thus improve the accuracy of character detection.

[0070] The area to be recognized for text is cut into equal - sized parts. The OCR model is used to detect characters in the pictures after being cut into a fixed size. Character cutting is performed according to the coordinates obtained from character detection and named automatically in the cutting order. Then, individual character recognition is performed on each cut character picture. By reducing other interference factors (cutting the character to be recognized from the original picture to reduce other content on the picture, such as lines and components that interfere with character recognition), the influence of other factors on the character recognition result is reduced. Finally, the built - in recognition accuracy of OCR is used to screen the recognition results. When the recognition accuracy is less than the threshold value, it is considered that the recognition is incorrect, and the character recognition result is deleted from the recognition table. When the recognition accuracy is greater than the threshold value, it is considered that the character recognition result is correct, and the result is retained. Among them, the structure of the recognition table is character corresponding to picture number, recognition result, the upper - left coordinate of the character in the original picture, and the lower - right coordinate of the character in the original picture. If it is retained, the whole data is retained; if it is deleted, the whole data is deleted.

[0071] The present invention improves the character recognition accuracy through the following technical solutions:

[0072] (1) Image cutting process. To achieve the purpose of reducing the character recognition area, it is necessary to use an OCR model to detect characters in the image. However, the result of character detection is related to the size of the image entering the model, so the image needs to be cut into a unified size. To obtain better character detection results, the image to be recognized needs to be cut into a fixed size. When the pixel size of the image is set between 550 and 850 (a square with a side length of 550 to 850 pixels), the best detection effect can be obtained. According to the set pixels, the image is cut into equal sizes using OpenCV, and the cut images are saved for passing into the OCR model for character detection.

[0073] (2) Character detection. The cut image is passed into the OCR model for character detection, and the characters are sheared according to the coordinate results detected by the model. Since OpenCV can shear according to pixel points, the detection coordinates are also set as pixel point coordinates, so that shearing can be performed according to the coordinates. The sheared character images are saved, and the saved results are reserved for subsequent steps of character recognition.

[0074] (3) Character recognition. The sheared character images are batch-passed into the OCR model for character recognition, and the recognition results and recognition accuracies are stored in the outer cut rectangle coordinate table (this table is automatically created according to the detection results of the characters after character detection, and stores the character numbers, the upper left coordinates of the characters, and the lower right coordinates of the characters). And the recognition results are screened according to the accuracy. When the accuracy is lower than the threshold, the recognition results are deleted, while those with a recognition accuracy higher than the threshold are retained. The threshold needs to be set above 80 to play a screening role.

[0075] (4) Character region coordinate conversion. Since the whole image is cut into multiple equal-sized regions, the coordinates of each character region are the coordinates corresponding to each small image. The coordinates need to be converted. According to the size of the cut image, the length of the cut image set by the multiple is added to the abscissa and the width of the cut image set by the multiple is added to the ordinate in turn, so as to obtain the true corresponding coordinates of the character region on the whole image.

[0076] When the present invention performs line recognition, it is based on the edge of the line for line recognition. The recognition line generated according to the edge of the line sandwiches the line in the middle, and the width of the line can be calculated only according to the edge coordinates of the line.

[0077] The present invention performs line recognition and obtains the line width through the following technical solutions:

[0078] (1)Character content and component content masking. For line detection in images, it is necessary to mask the character part and the component part in the image to reduce the interference of the lines in the characters and components on our recognition results.

[0079] (2)Line detection. On the basis of masking the characters and components in the image, use fast line detection based on OpenCV to obtain line data. Create a fast line detector class (this detector can be created by the createFastLineDetector function in OpenCV), and use the detector class to obtain the line vectors in the image. In this part, it is necessary to traverse the line vectors to obtain the coordinate information of the lines and save the x and y values of the starting coordinates and the x and y values of the ending coordinates of the lines. When saving, pay attention to the coordinate order of the lines. Determine the starting coordinate x value as the x value with a smaller value, and the ending coordinate x value as the x value with a larger value. Modify the y value according to the corresponding relationship.

[0080] (3)Line information screening and classification. According to coordinate operations, obtain the offsets of the x and y values; determine the horizontal and vertical relationships of the lines based on the obtained offsets. If the x coordinate offset is greater than the threshold and the y coordinate offset is less than the threshold, the line is a horizontal line; if the y coordinate offset is greater than the threshold and the x coordinate offset is less than the threshold, the line is a vertical line; if the offsets of both the x coordinate and the y coordinate are greater than the threshold, the line is a slant line. Store the classified line data separately. The threshold value should be between 5 and 30. If the threshold is less than 5, it may lead to unclear line classification, and if it is greater than 30, it may also affect the classification result.

[0081] (4)Line connection and merging processing. Process each line storage table (this line storage table is automatically created according to the classification during line information screening and classification, and stores the coordinate information of the lines and the x and y value offsets of the lines) respectively, connect and merge the broken short lines to make them into a longer line. When connecting and merging the horizontal lines, it is necessary to process according to their x-axis coordinates; when connecting and merging the vertical lines, it is necessary to process according to their y-axis coordinates. Since the slant lines in the power CAD drawings are generally short slant lines, no connection and merging operations are required for the slant lines.

[0082] (5) Line data clustering. Since the stored line data are all data obtained from line detection, we need to perform clustering operations on the horizontal lines, vertical lines, and diagonal lines after connection processing respectively to obtain the true data of the original lines and the widths of the lines to be recognized. When performing line clustering, set the threshold according to the known line width. This threshold is set according to the maximum line width. If standard line width pictures are used, setting the threshold between 10 and 20 can obtain better clustering results. Vertical lines are clustered according to the x coordinate, and the difference between the x coordinates of two lines is calculated to determine the line width. Horizontal lines are clustered according to the y coordinate, and the difference between the two y values is calculated to determine the line width. When clustering diagonal lines, it is necessary to calculate the center point coordinates of the lines and use these coordinates for clustering and line width calculation (the line width is equal to the absolute value of the difference between the two centroids). Among them, when clustering, the distance between the centroids of the lines is used as the judgment basis for clustering. When the distance between the centroids of the lines is less than the threshold, the two lines are considered to be the two sides of one line, that is, they are clustered into one line, and the distance between the centroids is the width of the line after clustering. The threshold is set to 10 - 20.

[0083] The steps of the method of the present invention are as follows:

[0084] (1) Cut the picture to be recognized into pictures of equal size;

[0085] (2) Sequentially input the cut pictures into the OCR model for character detection;

[0086] (3) Crop according to the coordinates of the character detection results, and cut out each character picture;

[0087] (4) Input the cropped character pictures into the OCR model for recognition, and screen according to the recognition accuracy rate to obtain the recognition results; the recognition results include character regions and character contents;

[0088] (5) Perform coordinate transformation on the character regions to obtain the true corresponding coordinates of the character regions on the entire original picture;

[0089] (6) Cover the characters and components on the original picture according to the character recognition results;

[0090] (7) Perform line detection on the covered picture;

[0091] (8) Classify the detected lines into vertical lines, horizontal lines, and diagonal lines;

[0092] (9) Connect and merge the vertical lines and horizontal lines respectively;

[0093] (10) Perform clustering on these three types of lines, namely vertical lines, horizontal lines, and diagonal lines, to obtain the coordinates of the lines and the corresponding line widths.

[0094] Application Example

[0095] Cut the electrical CAD drawing into equal-sized pictures. Input one of the pictures into the OCR model for character detection, and label the detection results to obtain the content as shown in Figure 5 the following.

[0096] Cut the characters in the picture according to the character detection results to obtain a series of pictures as shown in Figure 6 the following;

[0097] Send the cropped character pictures into the OCR model for recognition, and screen according to the recognition accuracy rate to obtain the recognition results; Figure 7 is the annotation of the recognition results on the original picture;

[0098] Cover the characters on the original picture according to the character recognition results, and cover the components to obtain the picture as shown in Figure 8 the following;

[0099] Perform line detection on the covered picture, classify the detected lines into vertical lines, horizontal lines, and oblique lines, obtain the corresponding line information table, and perform clustering on these three types of lines, namely vertical lines, horizontal lines, and oblique lines, to obtain the coordinates of the lines and the corresponding line widths. The line clustering result table is as shown in Figure 9 the following; Figure 9 In the above, x0 is the x coordinate of the starting point of the line, y0 is the y coordinate of the starting point of the line, x1 is the x coordinate of the end point of the line, y1 is the y coordinate of the end point of the line, and w is the width of the line.

[0100] The annotation of the line recognition results on the original picture is as shown in Figure 10 the following.

[0101] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for recognizing characters and lines with width in scanned electrical CAD drawings, characterized in that It includes the following steps: Perform equal-sized image cutting on the image to be recognized; Sequentially input the cut images into the OCR model for character detection; Crop according to the coordinate results of character detection, and cut out individual character images; Send the cut character images into the OCR model for recognition, and screen according to the recognition accuracy rate to obtain the recognition results; the recognition results include character regions and character contents; Perform coordinate transformation on the character regions to obtain the true corresponding coordinates of the character regions on the entire original image; Cover the characters on the original image according to the character recognition results, and cover the components; Perform line detection on the covered image to obtain line data, including the x and y values of the starting coordinates of the line and the x and y values of the ending coordinates; Classify the detected lines, and classify the lines into vertical lines, horizontal lines, and diagonal lines; Connect and merge the vertical lines and horizontal lines respectively; Cluster the three types of lines, namely vertical lines, horizontal lines, and diagonal lines, to obtain the coordinates of the lines and the corresponding line widths; 2. The method for identifying characters and line segments with widths in the scanned image of an electrical CAD drawing according to claim 1, wherein: Perform equal-sized image cutting on the image to be recognized, and cut it into squares with side lengths of 550 to 850 pixels; 3. The method for identifying characters and line widths of electrical CAD drawing scanned images according to claim 1, characterized in that: Send the cut character images into the OCR model for recognition. When the accuracy rate is lower than the threshold, delete the recognition results; when the recognition accuracy rate is higher than the threshold, retain the recognition results; The threshold is not less than 80%; 4. The method for identifying characters and lines with width in the scanned image of an electrical CAD drawing according to claim 1, wherein: The covered image uses fast line detection based on OpenCV to obtain line data; 5. The method for identifying characters and lines with width in the scanned image of the electrical CAD drawing according to claim 4, wherein: Determine the starting coordinate x value as the smaller value, and determine the ending coordinate x value as the larger value, and modify the y value according to the corresponding relationship; 6. The method for identifying characters and lines with width in the scanned image of an electrical CAD drawing according to claim 1, characterized in that: When classifying the detected lines, first calculate the offset of the x and y values; subtract the starting coordinate x value of the line from the ending coordinate x value of the line to obtain the x value offset, and subtract the starting coordinate y value from the ending coordinate y value of the line to obtain the y value offset; If the x coordinate offset is greater than the threshold and the y coordinate offset is less than the threshold, the line is a horizontal line; If the y coordinate offset is greater than the threshold and the x coordinate offset is less than the threshold, the line is a vertical line; If the offsets of both the x coordinate and the y coordinate are greater than the threshold, the line is a diagonal line; store the classified line data separately; 7. The method for identifying characters and line widths of the scanned image of an electrical CAD drawing according to claim 6, characterized in that: The threshold value ranges from 5 to 30; 8. The method for identifying characters and line widths of scanned electrical CAD drawings according to claim 1, wherein: Cluster the three types of lines, namely vertical lines, horizontal lines, and diagonal lines, to obtain the coordinates of the lines and the corresponding line widths; the specific method is: The vertical lines are clustered according to the x coordinates, and the difference between the x coordinates of the two lines is calculated to determine the line width. The horizontal lines are clustered according to the y coordinates, and the difference between the two y values is calculated to determine the line width. When clustering the diagonal lines, calculate the center point coordinates of the line, and use this coordinate for clustering and line width calculation; 9. The method for identifying characters and line widths of scanned images of electrical CAD drawings according to claim 8, characterized in that: When clustering, use the distance between the centroids of the lines as the judgment basis for clustering. When the distance between the centroids of the lines is less than the threshold, recognize the two lines as the two sides of one line, that is, cluster them into one line, and the distance between the centroids is the width of the clustered line. The threshold is set to 10 to 20.

Citation Information

Patent Citations

  • Test paper information extraction method and system and computer readable storage medium

    CN110414529A

  • Method and device for extracting lines in image, storage medium and electronic device

    CN110738219A