Engineering drawing calculation quantity extraction system and method
By constructing a coarse-grained-refined funnel-shaped recognition architecture and a five-level progressive row and column matching mechanism, the problem of character and table recognition in engineering drawings was solved, and efficient automatic extraction and accurate calculation of engineering quantities were achieved.
Patent Information
- Application Number
- CN202511501771.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
AI Technical Summary
Existing optical character recognition and table recognition technologies are difficult to apply to the recognition of material usage tables and the extraction of quantities in engineering drawings. They suffer from problems such as insufficient drawing resolution, complex character headers, and non-standard table rows and columns, resulting in low recognition efficiency and a high susceptibility to errors.
It adopts a coarse-grained-refined funnel-shaped recognition architecture, combined with a five-level progressive row and column matching mechanism, and automatically recognizes characters and tables in engineering drawings through OCR models and table recognition models to generate structured engineering quantity data.
It achieves efficient and automatic extraction of engineering quantities, improves recognition accuracy and precision, avoids errors in manual operation, and enhances the efficiency and automation of engineering quantity calculation.
Smart Images

Figure CN120976959A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to a system and method for extracting quantities from engineering drawings. Background Technology
[0002] Accurate calculation of quantities for steel reinforcement, concrete, and other materials in engineering projects is a core aspect of project budgeting and cost control. However, engineering drawings are mostly paper-based or scanned copies, lacking structured data and thus unable to be directly used for information-based calculations. Traditional methods rely on manually consulting material usage tables in engineering drawings and adding them up one by one to arrive at the result, which is inefficient and prone to errors.
[0003] With the in-depth development of artificial intelligence technology, Optical Character Recognition (OCR) and table recognition technologies have shown broad application prospects in the field of automated processing of engineering drawing information. OCR technology focuses on accurately converting text information in images into editable and computable text; table recognition technology aims to parse the table structure in images or documents and convert it into structured data. The core of these two technologies lies in accurately extracting target content from complex image information, making automated processing of engineering drawing information possible.
[0004] While existing optical character recognition (OCR) and table recognition frameworks perform well for recognizing ordinary tables, they still have limitations when facing the specific needs of the engineering industry, making them difficult to directly apply to the recognition of material usage tables and the extraction of quantities from engineering drawings. Specifically, engineering drawings present the following common challenges during the recognition process:
[0005] (1) Insufficient drawing resolution: Scanned or old drawings have low clarity, which affects the identification of details.
[0006] (2) Complex character headers: containing a large number of engineering symbols (such as “Φ” representing diameter) and multi-level composite headers.
[0007] (3) Non-standard table rows and columns: Engineering drawing tables often have merged cells, misaligned rows and columns, missing or discontinuous borders, etc., and the structure is complex.
[0008] While optical character recognition (OCR) and table recognition technologies have mature applications in other document processing fields, continuous technical optimization and solution improvement are still needed to meet the complex requirements of quantity surveying and recognition in engineering drawings, especially to meet industry standards for high accuracy and robustness. It is worth noting that the engineering field has accumulated massive amounts of historical drawing data, which provides valuable training resources for industry-specific model adaptation. Through targeted model training and functional development, it is expected to significantly improve the algorithm's adaptability to engineering scenarios and its recognition accuracy.
[0009] Given the characteristics of engineering drawings, how to achieve accurate recognition of material usage tables and automatic extraction of key field information and quantities through in-depth optimization of optical character recognition and table recognition technologies, and ultimately achieve efficient and accurate calculation of quantities, has become a pressing technical problem in the current engineering field. Summary of the Invention
[0010] In view of this, the purpose of this invention is to provide an engineering drawing quantity extraction system and method, especially an engineering drawing quantity extraction system and method based on optical character recognition and table recognition. By automatically recognizing the quantity tables in the drawings and extracting key information, the drawing information is intelligently converted into structured data that can be directly used for calculation and analysis, thereby achieving efficient automatic extraction of engineering quantities and solving the problems of low informatization level of current engineering drawings, low efficiency of manual quantity calculation and easy error.
[0011] To achieve the above objectives, the present invention adopts the following technical solution:
[0012] The first aspect of this invention proposes an engineering drawing quantity extraction system, comprising:
[0013] The input / output interface module includes an input interface unit and an image preprocessing unit. The input interface unit is used to receive and parse engineering drawings uploaded by users, and the image preprocessing unit is used to perform standardized preprocessing on the drawing images.
[0014] The OCR recognition module adopts a coarse-fine funnel-shaped recognition architecture consisting of a coarse-grained OCR model and a fine-grained OCR model. The coarse-grained OCR model is used to detect and recognize regular characters in text boxes, while marking engineering characters. The fine-grained OCR model is used to accurately recognize the marked engineering characters and correct text with specific character combinations.
[0015] The table recognition module includes a cell extraction model and a table row and column matching model. The cell extraction model is used to detect cells in engineering drawings, and the table row and column matching model is used to assign row and column IDs to cells.
[0016] The text and cell matching module is used to merge text boxes and match the text content to the corresponding cells to generate an Excel spreadsheet data file;
[0017] The quantity calculation semantic extraction module is used to convert Excel spreadsheet data files into structured JSON files for standard engineering quantity calculation extraction.
[0018] Preferably, the image preprocessing unit performs standardization preprocessing on the drawing image, including resolution enhancement processing, edge enhancement processing, image format conversion processing, and binarization processing, to output image data that meets the requirements of the recognition task.
[0019] Preferably, the coarse-grained OCR model is a lightweight detection-recognition joint model built on the PaddleOCR framework, which is pre-trained and optimized using a dataset containing engineering drawing data.
[0020] Preferably, the refined OCR model includes a geometric feature classification engine for accurate recognition of engineered characters marked by the coarse-grained OCR model. The workflow of the geometric feature classification engine is as follows:
[0021] (1) Extract binary images for engineering character feature analysis by connected component detection;
[0022] (2) Count the number of effective connected components in the image;
[0023] (3) Calculate the ratio of black pixels at the bottom and top of the image region;
[0024] (4) Based on the extracted binary image, effective connected component count and black pixel ratio index, the engineering characters are accurately identified using the preset decision rules.
[0025] Preferably, the refined OCR model includes a rule engine for correcting text with specific character combinations, including "engineering + number" characters, "Chinese characters + number" characters, and "pure number" characters. The correction by the rule engine includes:
[0026] (1) For the characters “engineering + number”, a geometric feature classification engine is used to establish an engineering symbol feature library, and the initial symbol is identified first and the subsequent characters are recursively verified.
[0027] (2) For the "Chinese character + number" character, use regular expressions to validate the "Chinese character + number" structure template;
[0028] (3) For “purely numerical” characters, corrections are made by using geometric information based on the easily identifiable structure and type of errors in engineering drawings.
[0029] Preferably, the table recognition module further includes a text pixel removal unit, which is used to remove pixels within the text box using the text box position information detected by OCR before cell extraction.
[0030] Preferably, the table row and column matching model uses a five-level progressive row and column matching mechanism to assign row and column IDs to cells, including:
[0031] (1) Initial geometric assignment: A primary mapping is established by calculating the intersection-union ratio of the text box and the cell;
[0032] (2) Multi-row merging optimization: Merging adjacent physical rows based on dynamic thresholds;
[0033] (3) Secondary allocation of overlapping areas: The Hungarian algorithm is used to perform optimal matching of conflict areas;
[0034] (4) Heuristic correction: Correcting unassigned text boxes based on Delaunay triangulation and conditional random field model;
[0035] (5) Cross-row and column detection: Identify merged cells that span columns, rows, or rows and columns and assign row and column information.
[0036] Preferably, the text and cell matching module merges the text boxes based on the text box coordinates and text box text information output by the OCR recognition module, and the cell coordinates and cell row and column ID information output by the table recognition module, by constructing an index space and geometric analysis, and then matches the merged text box text information to the corresponding cell.
[0037] Preferably, the quantity calculation semantic extraction module uses a dynamic header positioning mechanism and field mapping engine to convert Excel spreadsheet data files into structured JSON files for standard engineering quantity calculation extraction.
[0038] In another aspect, this invention also proposes a method for extracting quantities from engineering drawings, which employs the aforementioned engineering drawing quantity extraction system and includes the following steps:
[0039] Engineering drawing preprocessing: Receive and parse engineering drawings uploaded by users, and perform standardized preprocessing on the drawing images;
[0040] OCR Recognition: A coarse-grained OCR model is used to detect text boxes in the pre-processed engineering drawings and recognize regular characters in the text, while marking engineering characters; a refined OCR model is used to further accurately recognize the marked engineering characters and correct specific character combinations in the text.
[0041] Table recognition: Extract table cells and assign row and column IDs to cells using a five-level progressive row and column matching mechanism;
[0042] Text and cell matching: Based on the text box coordinates and text information detected by OCR, and the cell coordinates and cell row and column ID information provided by table recognition, the text boxes are merged, and the merged text box text information is matched to the corresponding cells to generate an Excel table data file;
[0043] Semantic extraction of quantity calculation: Convert Excel spreadsheet data into standard engineering quantity calculation data and output a structured JSON file.
[0044] The beneficial effects of this invention are as follows:
[0045] (1) The engineering drawing quantity extraction system provided by this invention takes the optimized OCR recognition model and table recognition model as its core and innovatively constructs a five-level cascaded recognition architecture to realize the end-to-end fully automatic extraction of engineering quantity information. Through key technologies such as coarse-grained-refined funnel-shaped recognition architecture, five-level progressive row and column matching mechanism, dynamic adaptive matching mechanism, and automated extraction of quantity calculation recognition, this system effectively solves the technical problems commonly found in engineering drawings, such as character adhesion, table deformation, and scanning noise, and significantly improves the efficiency and accuracy of engineering drawing quantity extraction.
[0046] (2) This invention addresses the problems of low character resolution, high scanning noise, and difficulty in recognizing engineering characters in engineering drawings by designing a coarse-grained-refined funnel-shaped recognition model. The coarse-grained OCR model quickly and accurately recognizes common Chinese and English characters, numbers, punctuation marks, etc., while the refined OCR model uses a geometric feature classification engine to flexibly and accurately identify uncommon and difficult-to-recognize engineering characters. Furthermore, a rule engine is used to correct specific character combinations in the text, improving the speed and accuracy of OCR recognition. In addition, the funnel-shaped design decouples the recognition of engineering characters from that of common characters. Newly added engineering character types do not require retraining the model; only a recognition strategy needs to be added to the refined recognition model. This solves the drawbacks of time-consuming and resource-intensive fine-tuning training for engineering characters, greatly increasing the flexibility and scalability of the character recognition model.
[0047] (3) This invention cleverly combines upstream and downstream information of the OCR recognition model and the table recognition model by constructing a five-level cascaded recognition architecture, which solves the difficult task of table recognition and character recognition of engineering drawing data, and realizes the accurate allocation of character information to the table. In the table recognition process, the position of the text box detected by OCR is innovatively provided to the table recognition model. By removing the corresponding pixels of the text box, the interference on table recognition is significantly reduced, thereby improving the accuracy and processing speed of table recognition.
[0048] (4) The quantity extraction module of the present invention automatically identifies and extracts relevant fields. Through a dynamic adaptive matching mechanism and field mapping engine, it introduces prior experience to correct errors in identification and common field mapping relationships, realizing a fully automated process of engineering quantity from identification, input and calculation, avoiding errors in manual operation, and improving the work efficiency and fully automated process of engineering quantity calculation.
[0049] (5) This invention performs denoising, upsampling, and binarization on the input data through image preprocessing, further improving the recognition accuracy and robustness of the system. Simultaneously, this invention specifically establishes an engineering drawing dataset. All data comes from drawing data in specific engineering projects and has undergone rigorous screening and cleaning. This dataset covers real data from the engineering field and scanning scenarios, and includes first, second, third, and fourth grade steel reinforcement characters from engineering text, providing high-quality data support for model training and testing of tasks such as OCR and table recognition, effectively ensuring the robustness and scalability of OCR and table recognition tasks. Attached Figure Description
[0050] Figure 1 This is a general framework diagram of an embodiment of the present invention;
[0051] Figure 2 This is a schematic diagram of the composition of the engineering drawing quantity extraction system according to an embodiment of the present invention;
[0052] Figure 3 This is a flowchart illustrating the engineering drawing quantity extraction method according to an embodiment of the present invention;
[0053] Figure 4 The image shown is a standardized preprocessed engineering drawing image in this embodiment of the invention.
[0054] Figure 5 This is a schematic diagram of the OCR recognition module process according to an embodiment of the present invention;
[0055] Figure 6 This is a schematic diagram of the table recognition module in an embodiment of the present invention;
[0056] Figure 7 This is a schematic diagram of the upstream data flow of the text and cell matching module in an embodiment of the present invention. Detailed Implementation
[0057] To make the objectives, techniques, and advantages of this invention clearer, the invention will be described clearly and completely below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0058] Terminology Explanation:
[0059] Optical Character Recognition (OCR) is a technology that converts scanned images or documents containing text into machine-readable text data. It identifies and encodes characters (Chinese and English, numbers, and symbols) by analyzing pixel patterns in the image, enabling computers to process, edit, and search this text information.
[0060] Table recognition is a document intelligent recognition technology that automatically locates table cell areas, identifies their row and column structure, and extracts the text and data content contained within them by analyzing document images. Its goal is to accurately convert tabular information in physical or digital images into a structured, editable, and analyzable data format.
[0061] Engineering drawing quantity extraction: This is an intelligent recognition technology applied in the engineering field. It automatically identifies and analyzes graphic elements, symbols, and dimensional information on engineering drawings (such as architectural, structural, and mechanical and electrical drawings) to extract key component parameters and match them with engineering quantity calculation rules. The goal is to achieve efficient and accurate automatic calculation of engineering quantities.
[0062] This invention provides a system and method for extracting quantities from engineering drawings, particularly a system and method based on optical character recognition (OCR) and table recognition. By deeply optimizing existing, high-performing OCR models and table recognition frameworks, the system and method can accurately identify text content containing engineering characters and precisely locate table areas and parse complex row and column structures (including merged cells and multi-level headers). Furthermore, as... Figure 1 As shown, this embodiment of the invention uses a trained and optimized OCR model and a table recognition model as the core. By constructing a five-level cascaded recognition architecture system, it sequentially performs operations such as coarse-grained-refined funnel-shaped optical character recognition, table recognition, text and cell matching, and quantity semantic extraction. Finally, it outputs structured engineering quantity data (such as the weight of steel bars at all levels, concrete volume, etc.) files for subsequent calculations or system integration, realizing end-to-end engineering quantity information extraction.
[0063] Specifically, see Figure 2 The present invention provides an engineering drawing quantity extraction system, which includes the following modules:
[0064] I. Input / output interface module, including input interface unit and image preprocessing unit. The input interface unit is used to receive and parse engineering drawings uploaded by users, and the image preprocessing unit is used to perform standardized preprocessing on the drawing images and output image data that meets the requirements of the recognition task.
[0065] II. The OCR recognition module adopts a coarse-fine funnel-shaped recognition architecture consisting of a coarse-grained OCR model and a fine-grained OCR model. The coarse-grained OCR model is used to detect and recognize regular characters in the text box, while marking engineering characters. The fine-grained OCR model is used to accurately recognize the marked engineering characters and correct text with specific character combinations.
[0066] III. Table recognition module, including cell extraction model and table row and column matching model. The cell extraction model is used to detect cells in engineering drawings, and the table row and column matching model is used to assign row and column IDs to cells.
[0067] IV. Text and Cell Matching Module: This module merges text boxes and matches the text content to the corresponding cells to generate an Excel spreadsheet data file.
[0068] V. Quantity Calculation Semantic Extraction Module: This module converts Excel spreadsheet data files into structured JSON files for standard engineering quantity calculation extraction.
[0069] Specifically, see Figure 3 The present invention provides a method for extracting quantities from engineering drawings, comprising the following steps:
[0070] S1. Engineering drawing preprocessing steps: Receive and parse the engineering drawings uploaded by the user, and perform standardized preprocessing on the drawing images;
[0071] S2, OCR recognition steps: The coarse-grained OCR model is used to detect text boxes in engineering drawings and recognize regular characters in the text, while marking engineering characters; the refined OCR model is used to further accurately recognize the marked engineering characters and correct specific character combinations in the text.
[0072] S3. Table recognition steps: Extract table cells and assign row and column IDs to cells using a five-level progressive row and column matching mechanism;
[0073] S4. Text and cell matching steps: Based on the text box coordinates and text information detected by OCR, and the cell coordinates and cell row and column ID information provided by table recognition, the text boxes are merged, and the merged text box text information is matched to the corresponding cells to generate an Excel table data file.
[0074] S5. Semantic extraction steps for quantity calculation: Dynamically locate the table header and map the project fields, convert the Excel table data into standard project quantity calculation data, and output a structured JSON file.
[0075] Furthermore, the method for extracting quantities from engineering drawings also includes:
[0076] S0. Model Construction and Training Optimization Steps: Construct and train the OCR recognition model and table recognition model. The OCR recognition model includes a coarse-grained OCR model and a fine-grained OCR model, while the table recognition model includes a cell extraction model and a table row and column matching model. Specifically, the coarse-grained OCR model is trained and optimized using a dataset containing engineering drawing data; the fine-grained OCR model is optimized using a geometric feature classification engine; and the table row and column matching model undergoes five levels of progressive row and column matching optimization.
[0077] The design structure and workflow of each module will be explained in further detail below.
[0078] I. Input / Output Interface Module
[0079] The input / output interface module is deployed to receive and parse engineering drawings uploaded by users, and performs standardized preprocessing on the drawings to provide high-quality image data input for subsequent OCR and table recognition tasks. This module includes an input interface unit and an image preprocessing unit. The input interface unit receives engineering drawings uploaded by users through a standardized interface, supporting multiple common image formats such as JPG, PNG, and PDF. Because the quality of the raw data uploaded by users is inconsistent (e.g., varying resolutions, noise), further preprocessing tasks are performed by the image preprocessing unit to denoise and upsample the images, ultimately outputting preprocessed image data that meets the requirements of subsequent recognition tasks, ensuring the accuracy and efficiency of the subsequent recognition process. The specific workflow within this module is as follows:
[0080] 1. The input interface unit receives and parses the engineering drawing files uploaded by the user.
[0081] The input interface unit uses a standardized input interface, such as HTTP, to establish a communication connection between the user end and the system end. Users can upload engineering drawing files to the system through this interface. After the user uploads the engineering drawing files, the system first checks them to ensure that the file upload address is correct, the format meets the requirements, and the file size is compliant. Then, the compliant files are parsed and converted into a standardized image format that the system can process, such as PIL image format, providing input data for the subsequent image preprocessing unit. It should be noted that this unit integrates a batch processing queue management mechanism, which can process dozens of drawings asynchronously for checking and parsing in parallel, avoiding the waiting delays of single-task processing and significantly improving the system's processing efficiency.
[0082] Furthermore, the input interface unit can automatically preload the trained and optimized OCR recognition model and table recognition model upon system startup. This ensures that model resources only need to be loaded once at system startup, avoiding the time loss caused by repeatedly loading models each time data is processed, thus achieving efficient processing. Since the model parameters are the system-trained and optimized model parameters, they are set to default values. Unless the user uploads new parameters, the system will use the predefined model parameters by default. It should be understood that users can also upload custom model parameters through the input interface to override the default values according to their actual needs.
[0083] 2. The image preprocessing unit further performs standardized preprocessing on the engineering drawings.
[0084] The image preprocessing unit takes the image data parsed by the input interface unit as input, employs a high-efficiency processing chain based on pixel rules, and achieves standardized processing of engineering drawings through multi-step joint debugging, ultimately outputting a high-quality image suitable for OCR recognition and table recognition. Processing includes:
[0085] (1) Resolution enhancement processing.
[0086] To address the potential issue of low resolution in the parsed image, a resolution enhancement operation is first performed. A bicubic interpolation algorithm is used to upsample the image by a factor of 2. This algorithm considers information from the surrounding 16 pixels during interpolation calculations, which can improve image resolution while better preserving image details and edge information. It avoids image blurring or jagged edges caused by simple interpolation algorithms, making it particularly suitable for processing complex lines and small text in engineering drawings.
[0087] (2) Edge enhancement processing.
[0088] Unsharpened mask filtering enhancement technology is applied to sharpen the edges of characters and table lines, further highlighting their edge features. This technique first blurs the image to obtain a mask image, then subtracts a certain proportion of the mask image from the original image, thereby enhancing areas with dramatic grayscale changes (i.e., edge areas). This makes character strokes and table lines clearer and more discernible, providing more explicit target features for subsequent recognition tasks.
[0089] (3) Image format conversion processing.
[0090] To speed up processing, the image is transformed twice:
[0091] a. First, convert the original PIL image format to a NumPy array so that subsequent image processing can be performed using function libraries. NumPy arrays are an efficient numerical computation data structure that can seamlessly interface with mainstream image processing libraries (such as OpenCV, Scikit-image, etc.).
[0092] b. After format conversion, convert the image to grayscale. This is done by calculating the weighted average of the RGB channels for each pixel (usually using the formula: Grayscale value = 0.299×R + 0.587×G + 0.114×B), converting the color image to a single-channel grayscale image. This step removes the interference of color information on text features, reduces the data dimensionality and complexity of the image, and preserves the grayscale difference between the text and the background.
[0093] (4) Binarization.
[0094] To maximize the contrast between text and background, the image preprocessing unit performs thresholding on the grayscale image to achieve binarization. Specifically, a fixed threshold method is used, with a threshold value of 220 and a maximum value of 255. When the grayscale value of a pixel in the grayscale image is greater than or equal to 220, the pixel value is set to 255 (white); when the grayscale value is less than 220, the pixel value is set to 0 (black). This operation clearly separates dark text and lines on a light background in engineering drawings, ultimately generating a clear binary image containing only black and white and fully preserving the original drawing's topological structure (such as line connections and character position distribution). This binary image, as the output of the image preprocessing unit, is directly used for subsequent OCR and table recognition tasks.
[0095] Figure 4 This shows an engineering drawing image that has been preprocessed by the input / output interface module.
[0096] II. OCR Recognition Module
[0097] The OCR recognition module adopts a coarse-grained-refined funnel-shaped recognition architecture, including a coarse-grained OCR model and a fine-grained OCR model. For example... Figure 4 As shown, engineering drawings contain common Chinese and English characters, numbers, punctuation marks, and also less common specialized engineering characters. For example... Figure 5 As shown, the OCR recognition module achieves efficient and accurate character recognition in engineering drawings through a two-level processing mechanism: First, a coarse-grained OCR model quickly and accurately detects the position of text boxes and recognizes most common characters, while marking more difficult-to-recognize engineering characters; then, a fine-grained OCR model accurately recognizes the engineering characters and corrects related text. The specific workflow is as follows:
[0098] 1. The coarse-grained OCR model recognizes regular characters while marking difficult-to-recognize engineering characters.
[0099] The coarse-grained OCR model is primarily responsible for the rapid detection and recognition of engineering drawings, performing initial recognition of easily identifiable characters while simultaneously filtering targets for finer processing. Specifically, it detects the positions of text boxes in engineering drawings, directly outputting the recognition results for common characters (Chinese and English, numbers, punctuation); for complex engineering characters, it performs a unified mapping and classification, such as mapping symbols for first-grade and second-grade rebar to the same category representing engineering characters (e.g., a unified output of [ENG_SYMBOL] or a special character like *), without fine differentiation. The final output includes the coordinates of the detected text boxes and the recognized text information, which may include confidence scores for individual characters in addition to the text characters themselves. For example, for... Figure 4 The model outputs the coordinates of the text box containing the character "Φ32": (x1, y1, x2, y2); the recognized text character is "[ENG_SYMBOL]32" and the confidence score for each character is [s1, s2, s3]. The label "[ENG_SYMBOL]" represents an engineering character and requires further refinement using a more precise OCR model.
[0100] Specifically, the coarse-grained OCR model is a lightweight detection-recognition joint model built on the PaddleOCR framework, employing a hybrid architecture of shared convolutional layers and independent long short-term memory (LSTM) artificial neural network heads. The shared convolutional layers efficiently extract low-level image features, reducing redundant computations; the independent LSM heads perform feature processing separately for detection and recognition tasks, improving task specificity.
[0101] The text detection module employs a differentiable binarized neural network structure, comprising a ResNet backbone network and an FPN feature pyramid. During operation, the engineering drawing image is first processed by the ResNet backbone network to extract multi-level feature maps, then input into the FPN feature pyramid for fusion. Finally, the differentiable binarized neural network structure predicts the text box position based on the fused feature maps, achieving accurate localization of the text region.
[0102] The text recognition module employs a convolutional recurrent neural network architecture, consisting of convolutional layers, recurrent layers, and a decoder. The convolutional layers use MobileNetV3 to extract character features, the recurrent layers use a bidirectional long short-term memory network to process sequence features, and the decoder uses a sequence data classification decoder to align character sequences, ultimately outputting the recognized text characters.
[0103] It should be noted that, to enhance the model's ability to recognize engineering characters and map all engineering characters to a single recognition category, thereby filtering out engineering characters and identifying specific types in the refined recognition stage, this embodiment of the invention pre-optimized and trained the coarse-grained OCR model using the open-source ICDAR2015 dataset and a self-built engineering drawing dataset. All data in the engineering drawing dataset comes from drawings in specific engineering projects, including engineering characters such as first, second, third, and fourth grade rebar characters. The data was filtered and cleaned to provide real-world data from the engineering field and scanning scenarios for model training. Furthermore, since engineering characters constitute a relatively small proportion, the engineering character data was resampled and its training order was shuffled during training. The trained model supports single and combined recognition of Chinese and English characters, numbers, punctuation, and engineering symbols, including over 6800 character sets, the ASCII character set, and created engineering character categories. The recognition accuracy for Chinese and English characters, numbers, and punctuation is the same as that of the open-source Paddle OCR, and the accuracy for engineering character recognition reaches 98.3%.
[0104] 2. The refined OCR model accurately recognizes engineering characters and corrects text with specific character combinations.
[0105] The refined OCR model addresses the challenges posed by the coarse-grained OCR model in recognizing engineering characters or specific character combinations. By triggering a geometric feature classification engine and a rule engine, it further refines and corrects these difficult-to-recognize characters, thereby improving the OCR model's accuracy. The model's output is the precisely recognized and corrected text content along with its corresponding text box positions, providing high-precision data for subsequent operations such as cell extraction and merging. Its specific workflow is as follows:
[0106] (1) Use a geometric feature classification engine to accurately identify engineering characters.
[0107] It should be noted that the refined OCR model has been pre-optimized for the geometric feature classification engine. The geometric feature classification engine, based on the geometric characteristics of the engineering characters selected by the coarse-grained OCR model, designs a recognition mechanism to classify the engineering characters. By deeply mining the geometric features of the engineering characters, it achieves accurate differentiation between different types of engineering characters. Its input includes the output of the engineering characters recognized by the coarse-grained OCR model, namely the position of the text box containing the engineering characters, text information, and confidence score recognized by the coarse-grained OCR model. Specific steps include:
[0108] a. Extract the target region.
[0109] Specifically, the position of the engineering characters is located in the text box containing the engineering characters identified by the coarse-grained OCR model. Then, a small image containing the engineering characters is cropped from the preprocessed engineering drawing image as the original cropped image. A connected component detection algorithm is used to extract all external connected components from the binarized original cropped image. A connected component is a region in an image consisting of adjacent pixels with the same pixel value.
[0110] The detection results are sorted according to the area of the connected components. The connected component with the largest area is selected as the target region, and the coordinate parameters of the boundary rectangle of the connected component are obtained, including the horizontal and vertical coordinates of the starting position, as well as the width and height of the region.
[0111] Based on the extracted boundary rectangle parameters, the original cropped image is cropped a second time to further extract the region of interest for feature analysis of the engineering characters, so as to remove more background and interference and focus on the engineering characters themselves.
[0112] b. Statistical analysis of effective connected components.
[0113] A counter is initialized using the effective connected component counting method to record the number of effective connected components.
[0114] The algorithm iterates through all detected connected components. For each component, its area is first calculated and checked against a threshold. Simultaneously, it checks if the component is located at an image edge. The area threshold needs to be adjusted based on common engineering character sizes, typically determined through extensive sample statistics to ensure that noisy connected components with excessively small areas are filtered out. Connected components located at image edges may be due to incomplete image cropping, not representing the complete engineering character, and therefore need to be excluded. Whether a component is located at an edge is determined by the distance between the boundary coordinates of the connected component and the image boundary. If the distance is less than a preset value, the connected component is considered to be located at an image edge.
[0115] If a connected component satisfies the area condition and is not an edge connected component, it is a valid connected component, and the counter is incremented by 1. The number of valid connected components is one of the important characteristics that distinguishes different engineering characters. For example, some engineering characters are composed of multiple parts, and the number of their valid connected components will differ from other characters.
[0116] c. Calculate the proportion of black pixels.
[0117] Calculate the black pixel ratio feature at the bottom and top of an image region. The bottom of the image region refers to the area extending upwards from the bottom of the region by a certain percentage (e.g., 10%-20%), and the top refers to the area extending downwards from the top of the region by a certain percentage (e.g., 10%-20%). The black pixel ratio is the ratio of the number of black pixels in this region to the total number of pixels in the region. This feature reflects the distribution of engineering characters at the top and bottom. Different engineering characters may show significant differences in this feature; for example, some engineering characters may have a wider top and a higher black pixel ratio, while others may have a wider bottom.
[0118] d. Classification of engineering characters.
[0119] The extracted binary image, effective connected component count, and calculated feature parameters are passed to the classification module. The binary image provides overall shape information of the engineering character, the effective connected component count reflects the number of components of the character, and feature parameters such as the proportion of black pixels describe the local features of the character from different perspectives. These geometric indicators together constitute the feature set for distinguishing engineering characters.
[0120] By integrating all feature information, pre-defined decision rules are applied to classify difficult-to-recognize engineering characters. These decision rules are formulated based on feature analysis of a large number of engineering character samples with known categories. For example, for Grade I and Grade II rebar symbols, corresponding judgment conditions are set by analyzing features such as the number of effective connected components and the proportion of black pixels. When the features of an input engineering character satisfy the decision rule corresponding to a certain category, it is classified into that category.
[0121] Finally, the specific category of the engineering characters is determined, and the original recognition results of the coarse-grained OCR model are returned and updated. The classification method using multi-geometric feature fusion integrates feature information from multiple dimensions. Compared to single-feature classification, it can more comprehensively and accurately describe the characteristics of engineering characters, significantly improving the accuracy and robustness of text recognition, especially engineering character recognition.
[0122] (2) Use a rule engine to correct text with specific character combinations.
[0123] After the geometric feature classification engine processes the engineering characters to obtain specific classification results, the rule engine is further used to correct the text of specific character combinations.
[0124] The rule engine mainly corrects three types of characters, including "engineering + number" characters, "Chinese character + number" characters, and "pure number" characters. These three types of characters are interfered by similar-shaped fonts during recognition, resulting in a high probability of misrecognition, and the confidence score is often lower than the threshold. For example, "均1024.4" is misrecognized as "均IO24.4", and "2×100 / 15" is misrecognized as "2xIOOII5". Through statistics, it is found that: a. When arranging characters on an engineering drawing, if only the first character is a Chinese character and the remaining characters are not Chinese characters, then all characters except the first character are numbers or decimal points; b. If the first and last characters are numbers, then the middle characters are also numbers or arithmetic symbols. Therefore, if the recognized text belongs to the above three types of specific character combination texts, the rule engine will be triggered. In some embodiments, to improve processing efficiency, when the recognized text belongs to the above three types of specific character combination texts and there are characters with a confidence score lower than the threshold, the rule engine is triggered.
[0125] For different types of characters, the rule engine will process them according to predefined prior rules. The prior rules are defined as the following 3 items: (1) If the first character is an engineering character, it is a "engineering + number" character combination, and the engineering character can only appear at the beginning of the text; (2) If there is and only the first character is a Chinese character, then the remaining characters should be numbers or decimal points; (3) When the first and last characters are numbers, the middle characters can only be numbers or arithmetic operators (including addition, subtraction, multiplication, division, decimal points, etc.).
[0126] Through the above rules, different treatments are carried out for different types of characters. For "engineering + number" characters, a geometric feature classification engine is used to establish an engineering symbol feature library, and after preferentially recognizing the starting symbol, the subsequent characters are recursively verified. For "Chinese character + number" characters, a regular expression is applied to verify the "Chinese character + number" structure template. For "pure number" characters, in view of the structures and types that are easily misrecognized in engineering drawings, geometric information is used for correction.
[0127] III. Table Recognition Module
[0128] The table recognition module includes a cell extraction model and a table row-column matching model. The cell extraction model detects the cells in the engineering drawing, and then the table row-column matching model assigns corresponding row IDs and column IDs to the cells. Its input is the preprocessed engineering drawing and the text box coordinate information detected by the OCR module, and the output is the recognized cell coordinates, cell row-column IDs and other information. The work process is as Figure 6 shown, specifically including:
[0129] 1. Remove the pixels within the text box.
[0130] Before extracting cells, you can use the "Text Pixel Clear Cells" function. This involves processing the engineering drawing data using the coordinates of the text boxes detected by OCR, removing pixels within the OCR text box area (i.e., setting them to white). After processing, the engineering drawing only retains the pixel information of the table. This reduces the impact of text on the table and speeds up table recognition.
[0131] 2. Use a cell extraction model to identify table cells.
[0132] The cell extraction model uses the Surya table recognition algorithm to extract cells. It employs an encoder-decoder architecture, with both working together to extract cells. The encoder uses a sliding window transformation layer to extract multi-scale features from the image, while the decoder uses an optimized transformation layer architecture for predicting the table structure.
[0133] Specifically, the encoder uses a sliding window transformation layer to extract image features. It segments the image into non-overlapping local patches, each of which is converted into a vector representation through an embedding layer. These vector representations are processed through multiple sliding window transformation layers to capture both local and global features of the image. The hierarchical design of the sliding window transformation layers enables the model to extract features at different scales, thus providing a better understanding of the image content.
[0134] The decoder predicts the table structure from the feature representation generated by the encoder. First, the decoder converts the table's attribute labels into embedding vectors through a label embedding layer. These embedding vectors are then fed into stacked decoder layers. Within each layer, the decoder combines image features provided by the encoder to perform feature transformation on the embedding vectors and progressively generates or refines the prediction sequence through a self-attention mechanism. Finally, the multi-layered representation is fed into multiple parallel linear layers, which output predictions for various table attributes, including bounding box regression coordinates and probability distributions for cell, row, and column categories.
[0135] 3. Use a table row and column matching model to match the row and column ID information of the cells.
[0136] Cell extraction models can accurately locate all potential table cells in an image, but the correspondence between these cells and their logical row and column IDs has not yet been established. The core task of table row and column matching models is to accurately map these physically detected cell bounding boxes to the logical row and column structure of the table, assigning the correct row and column IDs to each cell.
[0137] Because the table structures in actual engineering drawings are complex and varied, including scenarios with line breaks, overlaps, and merged cells, simple geometric matching methods are ineffective. Therefore, this model has undergone a five-level progressive row and column matching optimization beforehand.Figure 6 Through a step-by-step iteration from basic geometric relationships to intelligent correction, precise matching of cell and row / column IDs is achieved. The specific process is as follows:
[0138] (1) Initial geometric assignment.
[0139] A preliminary mapping relationship between cells and rows / columns is established, and the intersection-union ratio (IUR) of OCR text detection boxes and table cells is calculated. When the IUR is greater than a threshold, an association is established. Specifically, based on the cell coordinates output by the cell extraction model and the text detection box coordinates provided by the OCR module, for each OCR text detection box, all cells are traversed, and the IUR is calculated using the formula: (area of overlapping region) / (area of text detection box + area of cell - area of overlapping region). A threshold for the IUR is set (e.g., 0.7). When the IUR of an OCR text detection box is greater than this threshold, the text detection box is associated with the corresponding cell. At this time, the positional information carried by the text detection box can indirectly reflect the potential row / column position of its cell. For detection boxes with an IUR less than or equal to the threshold, they are marked as unassociated and left for subsequent processing. Through initial geometric allocation, text box-cell pairs with clear positional relationships are quickly identified, obviously mismatched combinations are filtered out, a preliminary mapping relationship between cells and rows / columns is established, and the initial outline of the table's row and column structure is sketched.
[0140] (2) Optimization of merging multiple lines.
[0141] In engineering drawings, a common problem is that line breaks in text can cause a single logical cell (containing multiple lines of text) to be incorrectly split into multiple adjacent physical rows. Multi-line merging optimization achieves line box fusion through dynamic threshold judgment. Specifically, in the preliminary row structure established in the previous step, adjacent rows in physical location are identified, and the vertical gap value between all adjacent row pairs is calculated. The median of all calculated line gap values is taken as the dynamic threshold. Adjacent row pairs are traversed; if the actual gap value between them is less than the dynamic threshold, it is considered that these two lines of text are likely to belong to the same logical cell. The bounding boxes of these two or more consecutive adjacent rows that meet the condition are geometrically merged to form a new, larger merged cell bounding box. Simultaneously, the row number topology is updated, the original rows to be merged are marked as merged, and a new, single row ID is used to represent this merged logical row, ensuring that subsequent matching is performed on the correct logical row structure.
[0142] (3) Secondary allocation of overlapping regions.
[0143] The initial filtering process addresses conflict regions, specifically all text boxes that remain unassigned after initial allocation and multi-line merging, along with all overlapping cells. These text box-cell candidate pairs constitute a set of conflicts to be resolved. Within this set, a local maximum intersection-union ratio (MOU) is calculated, and the Hungarian algorithm is used to achieve optimal matching, finding the most reasonable cell for the previously unassigned text. The weight function is as follows:
[0144] weight = 0.6 × intersection-union ratio + 0.2 × text similarity + 0.2 × centroid distance
[0145] Wherein, the intersection-union ratio is the intersection-union ratio of the text box and the cell, the text similarity is the similarity between the text content of the text box and the assigned text content of the cell, and the center point distance is the Euclidean distance from the center point of the text box to the center point of the cell.
[0146] (4) Heuristic correction.
[0147] After the first three steps, a relatively high proportion of text boxes remain unassigned. When this proportion exceeds a threshold, the heuristic layout module is activated. Specifically, it constructs a cell spatial distribution density map based on Delaunay triangulation. For each unassigned text box, its relative positional features with neighboring cells are extracted, and the horizontal relative offset rate is calculated. This is the horizontal offset of the text box's center point relative to the center points of neighboring cells, normalized to the cell width. A smaller horizontal relative offset rate indicates that the text box is more likely to belong to that cell's column. Similarly, the vertical offset features can be calculated. Then, the assignment probability is calculated using a conditional random field model, iteratively optimized until convergence. Based on the converged probabilities, the candidate cell with the highest probability is selected for assignment to each unassigned text box.
[0148] (5) Cross-row and column detection.
[0149] This step aims to accurately identify merged cells spanning columns, rows, or rows and columns within the table and assign them the correct row and column information. Specifically, it iterates through the cells row by row, checking if the right boundary of the current cell coincides with the left boundary of the next cell in the same row if the overlap exceeds a threshold. If it does, the two are merged. Similarly, it iterates through the cells column by column, checking if the bottom boundary of the current cell coincides with the top boundary of the next cell in the same column if the overlap exceeds a threshold. If it does, the cells are merged vertically. Then, the initially detected merged areas are traversed in reverse, checking for horizontal merges from right to left and vertical merges from bottom to top to verify continuity and avoid erroneous merges due to detection errors. This generates a cross-row and column topology map, ensuring the consistency of merged cells with the basic grid.
[0150] IV. Text and Cell Matching Module
[0151] The text and cell matching module aims to merge text boxes in a table using index space and geometric analysis, then accurately fill the text content into the corresponding table cells, ultimately generating a structured Excel-formatted table. For example... Figure 7 As shown, the input to this module includes the coordinates of the text box location and the text information recognized by the OCR recognition module, and the cell location coordinates and row / column IDs output by the table recognition module. Its specific workflow is as follows:
[0152] 1. Text box merging, including:
[0153] (1) Construct an index space for efficient querying and management of the spatial relationships of text boxes.
[0154] Specifically, the coordinate information of all text boxes detected by OCR is inserted into a preset index space so that other text boxes that are spatially adjacent or overlapping with a certain text box can be quickly queried in the future, avoiding brute-force traversal of all text boxes and improving the efficiency of subsequent steps.
[0155] (2) Merge vertical text boxes.
[0156] For each text box A in the spatial index:
[0157] i. Query and retrieve the candidate merge box B in the vertical direction of the current text box A by expanding the top and bottom boundaries of the current text box A.
[0158] ii. For each candidate merge box B, perform the following operation:
[0159] a. Table line detection
[0160] Detect whether there is a table line between text boxes A and B:
[0161] If table lines exist, skip candidate merge box B, assuming that the text boxes do not need to be merged, and re-enter step a to check the next candidate merge box B;
[0162] If no table lines exist, proceed to step b and perform an overlap ratio check.
[0163] b. Overlap ratio check
[0164] Calculate the overlap ratio between the current text box A and the candidate merge box B along the Y-axis. The formula for calculating the overlap ratio is:
[0165]
[0166] Where a1, a2 and b1, b2 are the boundaries of the two intervals on the Y-axis, and lengtha and length b These are the lengths of the two intervals on the Y-axis:
[0167] If the overlap ratio is less than the threshold, the two text boxes are considered not to overlap enough, the merging is skipped, and the process returns to step a to check the next candidate merge box B.
[0168] If the overlap ratio is greater than or equal to the threshold, proceed to the next step c: merging the text boxes.
[0169] c. Merge text boxes
[0170] If candidate merge box B passes the table line detection and overlap ratio check, it is merged with the current text box A. The boundary of the merged text box M is the extreme value of the two box boundaries:
[0171]
[0172] Among them, (x) a1 y a1 ), (x a2 y a2 These are the coordinates of the top-left and bottom-right corners of the first text box, respectively. b1 y b1 ), (x b2 y b2 The coordinates of the top-left and bottom-right corners of the second text box are shown below. To clarify the positional relationship, the coordinate system used has the top-left corner as the origin, the positive X-axis pointing horizontally to the right, and the positive Y-axis pointing vertically downwards. The merged text box M will replace the original two text boxes, and the candidate merged box B will be marked as merged.
[0173] iii. Continue checking steps (i)-(ii) for the merged text box M, attempting to merge it with more candidate boxes that meet the criteria. Continue until there are no new candidate merge boxes B that meet the criteria to merge.
[0174] (3) Merge horizontal text boxes.
[0175] Similarly, when performing a text box merge operation in the horizontal direction, the decision to merge is determined by expanding the left and right boundaries of the current text box. The process is completely symmetrical to the vertical merge in step 2, except that the operation direction is changed from vertical to horizontal, which will not be described further here.
[0176] 2. Match text to cells.
[0177] After text box merging, the scattered text lines are effectively aggregated into logically complete text blocks. Then, the cell geometric positions and row and column IDs output by the upstream table recognition module are spatially matched with the merged text content recognized by OCR, based on positional inclusion relationships or high intersection-union ratio.
[0178] Finally, the module fills the matched text content into its corresponding cells, generating a structured Excel spreadsheet data file as the final output.
[0179] V. Semantic Extraction Module for Quantity Calculation
[0180] The quantity calculation semantic extraction module is used to convert Excel data files into standard engineering quantity calculation data, and performs semantic mapping on the quantity calculation fields to be extracted. The input is the recognized Excel file, and the output is a structured JSON file containing the extracted quantity calculation data. Its workflow is as follows:
[0181] 1. Table header positioning.
[0182] Since the header position of engineering drawing tables is not fixed, the header positioning layer is determined first. The specific operation is as follows: dynamically scan the first 6 rows of the table, determine the header position based on the maximum non-empty column width criterion, calculate the effective text column width of each row, select the row with the largest column width as the candidate header row, and verify the keyword density of the header, where the keyword frequency is greater than 1.5 times the average vocabulary of the row.
[0183] 2. Semantic mapping.
[0184] After the table header is located, the semantic mapping stage begins. A field mapping engine maps heterogeneous header text in the table to predefined standard fields, resolving the issue of inconsistent engineering data formats. Specifically, the following steps are taken: Predefine engineering quantity calculation fields (such as "rebar diameter", "total weight", and "component length") and build a standard field library; for non-standard expressions that may appear in the header (such as "diameter" corresponding to "rebar diameter", and "length" corresponding to "component length"), a fuzzy matching system is built to automatically associate the closest standard field; through substring matching optimization, the longest common subsequence is extracted from heterogeneous headers to extract core substrings and accurately match standard fields.
[0185] 3. Data extraction and JSON generation.
[0186] After determining the standard fields through semantic mapping, the corresponding column attribute values in the Excel file can be extracted row by row to obtain the corresponding attribute information and generate a structured JSON file. The generated structured JSON file fully retains key engineering quantity attributes such as rebar number, diameter, number of rebars, and total weight, while also attaching metadata on the original coordinate location confidence level. This JSON file conforms to digital engineering delivery standards and can be directly imported into BIM quantity calculation software for engineering quantity verification.
[0187] This invention, through modular assembly line operations, automates the entire process from receiving original drawings, OCR recognition, table structure reconstruction, to standardized output. When processing engineering drawings, the collaborative work of the aforementioned modules generates a structured data file containing complete engineering quantity calculation information, including the location, specifications, quantity, and weight of reinforcing bars, meeting the needs of digital delivery in modern construction engineering.
Claims
1. A quantity extraction system for engineering drawings, characterized in that, include: The input / output interface module includes an input interface unit and an image preprocessing unit. The input interface unit is used to receive and parse engineering drawings uploaded by users, and the image preprocessing unit is used to perform standardized preprocessing on the drawing images. The OCR recognition module adopts a coarse-fine funnel-shaped recognition architecture consisting of a coarse-grained OCR model and a fine-grained OCR model. The coarse-grained OCR model is used to detect and recognize regular characters in text boxes, while marking engineering characters. The fine-grained OCR model is used to accurately recognize the marked engineering characters and correct text with specific character combinations. The table recognition module includes a cell extraction model and a table row and column matching model. The cell extraction model is used to detect cells in engineering drawings, and the table row and column matching model is used to assign row and column IDs to cells. The text and cell matching module is used to merge text boxes and match the text content to the corresponding cells to generate an Excel spreadsheet data file; The quantity calculation semantic extraction module is used to convert Excel spreadsheet data files into structured JSON files for standard engineering quantity calculation extraction.
2. The engineering drawing quantity extraction system according to claim 1, characterized in that, The image preprocessing unit performs standardized preprocessing on the drawing image, including resolution enhancement, edge enhancement, image format conversion, and binarization, to output image data that meets the requirements of the recognition task.
3. The engineering drawing quantity extraction system according to claim 1, characterized in that, The coarse-grained OCR model is a lightweight detection-recognition joint model built on the PaddleOCR framework. The model is pre-trained and optimized using a dataset containing engineering drawing data.
4. The engineering drawing quantity extraction system according to claim 1, characterized in that, The refined OCR model includes a geometric feature classification engine for accurate recognition of engineered characters marked by the coarse-grained OCR model. The workflow of the geometric feature classification engine is as follows: (1) Extract binary images for engineering character feature analysis by connected component detection; (2) Count the number of effective connected components in the image; (3) Calculate the ratio of black pixels at the bottom and top of the image region; (4) Based on the extracted binary image, effective connected component count and black pixel ratio index, the engineering characters are accurately identified using the preset decision rules.
5. The engineering drawing quantity extraction system according to claim 1, characterized in that, The refined OCR model includes a rule engine for correcting text with specific character combinations, including "engineering + number" characters, "Chinese characters + number" characters, and "pure number" characters. The correction by the rule engine includes: (1) For the characters "engineering + number", a geometric feature classification engine is used to establish an engineering symbol feature library, and the initial symbol is identified first and the subsequent characters are recursively verified. (2) For the "Chinese character + number" character, use regular expressions to validate the "Chinese character + number" structure template; (3) For “purely numerical” characters, corrections are made by using geometric information based on the easily identifiable structure and type of errors in the engineering drawings.
6. The engineering drawing quantity extraction system according to claim 1, characterized in that, The table recognition module also includes a text pixel removal unit, which is used to remove pixels in the text box using the text box position information detected by OCR before cell extraction.
7. The engineering drawing quantity extraction system according to claim 1, characterized in that, The table row and column matching model uses a five-level progressive row and column matching mechanism to assign row and column IDs to cells, including: (1) Initial geometric assignment: A primary mapping is established by calculating the intersection-union ratio of the text box and the cell; (2) Multi-row merging optimization: Merging adjacent physical rows based on dynamic thresholds; (3) Secondary allocation of overlapping areas: The Hungarian algorithm is used to perform optimal matching of conflict areas; (4) Heuristic correction: Correcting unassigned text boxes based on Delaunay triangulation and conditional random field model; (5) Cross-row and column detection: Identify merged cells that span columns, rows, or rows and columns and assign row and column information.
8. The engineering drawing quantity extraction system according to claim 1, characterized in that, The text and cell matching module, based on the text box coordinates and text information output by the OCR recognition module, and the cell coordinates and cell row and column ID information output by the table recognition module, merges the text boxes by constructing an index space and geometric analysis, and matches the merged text box text information to the corresponding cell.
9. The engineering drawing quantity extraction system according to claim 1, characterized in that, The quantity calculation semantic extraction module uses a dynamic header positioning mechanism and field mapping engine to convert Excel spreadsheet data files into structured JSON files for standard engineering quantity calculation extraction.
10. A method for extracting quantities from engineering drawings, characterized in that, The engineering drawing quantity extraction system according to any one of claims 1-9 includes the following steps: Engineering drawing preprocessing: Receive and parse engineering drawings uploaded by users, and perform standardized preprocessing on the drawing images; OCR Recognition: A coarse-grained OCR model is used to detect text boxes in the pre-processed engineering drawings and recognize regular characters in the text, while marking engineering characters; a refined OCR model is used to further accurately recognize the marked engineering characters and correct specific character combinations in the text. Table recognition: Extract table cells and assign row and column IDs to cells using a five-level progressive row and column matching mechanism; Text and cell matching: Based on the text box coordinates and text information detected by OCR, and the cell coordinates and cell row and column ID information provided by table recognition, the text boxes are merged, and the merged text box text information is matched to the corresponding cells to generate an Excel table data file; Semantic extraction of quantity calculation: Convert Excel spreadsheet data into standard engineering quantity calculation data and output a structured JSON file.
Citation Information
Patent Citations
English word identification method and apparatus
CN106127118A
Image key text extraction method and device
CN118887649A
Method and system for extracting structured data table of power grid engineering drawing and medium
CN119992577A
Complex large drawing table extraction and reconstruction method and system
CN120496108A
Recognition of printed characters using optical scanning - with obtained signals identified by neural network trained with signal patterns for different lines of character orientations
DE19507059A1
Cited By
Semantic sorting method and recognition system for hydraulic cylinder engineering drawing table
CN121963216A