Material information identification statistical method and device based on CAD export, electronic equipment, medium and program product

By parsing CAD drawings and filtering target layers based on rules, combining dual scripts to extract text and cell boundaries, establishing row and column index mapping, and introducing a deep table structure recognition model to complete missing boundary lines, the problem of locating and reconstructing material tables in CAD drawings is solved, and efficient and accurate material information statistics are achieved.

CN120764487AActive Publication Date: 2025-10-10ZHUHAI HUACHENG ELECTRIC POWER DESIGN INST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511197552.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-10
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to stably locate and reconstruct material tables in CAD-exported drawings, resulting in low material information statistical efficiency and high error rate. In particular, when the table boundary lines are missing or there is coordinate noise, the table structure cannot be effectively reconstructed.

Method used

By parsing DWG format engineering drawings, filtering target layers based on preset layer name matching rules, combining dual scripts to extract text and cell boundaries, establishing row and column index mapping, and introducing a deep table structure recognition model to complete missing boundary lines, a formatted Excel file is generated.

Benefits of technology

It achieves efficient, accurate and structured output of material information, significantly improves statistical efficiency and reduces the error rate of manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764487A_ABST
    Figure CN120764487A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and particularly discloses a material information recognition statistical method and device based on CAD export, electronic equipment, a medium and a program product.The method comprises the steps that a target layer is screened by analyzing a DWG drawing, a table object containing material information is automatically recognized, and characters and cell boundary coordinates of the table object are extracted; pairing the character file and the coordinate file based on the table ID, converting boundary coordinates into ordered row and column indexes, and establishing a mapping relation between characters and cells; and carrying out double sorting on the characters according to coordinates and splicing to generate complete cell contents. According to the method, accurate positioning, content restoration and structured output of the material table can be realized without manual intervention, the statistical efficiency is improved, and the error rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and specifically relates to a material information identification and statistics method, device, electronic equipment, medium and program product based on CAD export. Background Art

[0002] Engineering drawing material tables are often embedded in DWG drawings exported from CAD. Traditional statistical material tables rely on manual work or simple scripts / OCR, making it difficult to stably locate table objects in multiple layers and the sources are messy. In some existing technologies, even if text and boundary coordinates are obtained, it is difficult to establish a reliable mapping of text cells. In addition, when there are missing boundary lines or coordinate noise in the table, it is impossible to stably reconstruct the table structure for statistics. The above problems make it difficult to effectively reconstruct material tables from CAD-exported drawings and directly generate usable structured Excel files, which is inefficient and has a high error rate. Summary of the Invention

[0003] In this regard, the present invention provides a material information identification and statistical method, device, electronic device, medium and computer program product based on CAD export to solve the above technical problems.

[0004] The present invention provides a material information identification and statistical method based on CAD export, comprising the following steps:

[0005] Step S101: Parse a DWG format engineering drawing to obtain all layer sets in the drawing, traverse the layer set based on a preset layer name matching rule to filter and obtain a target layer, identify a table object from the target layer, and record the number of tables and the global coordinate range of the table for the table object, wherein the table object contains material information;

[0006] Step S102: for each identified table object, extract the text elements and their spatial coordinate information in the table using a first script to generate a text data file; and extract the cell boundary coordinates using a second script to generate a cell coordinate file;

[0007] Step S103, pairing the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and converting the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file;

[0008] Step S104: extracting text data and cell data from the paired file, wherein the text data includes each text element and its spatial coordinates, and the cell data includes cell boundary coordinates and their corresponding row and column indexes. An association is established between the text and the cell based on the coordinates of the text elements and the cell boundary coordinates, forming a mapping between the cell row and column indexes and the text.

[0009] Step S105 , based on the mapping between the cell row and column indexes and the text, double sorting the text coordinates, and concatenating the sorted text to obtain the complete text content of the corresponding cell.

[0010] In another aspect, the present application further provides a material information identification and statistics device based on CAD export, comprising:

[0011] A table parsing and identification module is used to parse a DWG format engineering drawing to obtain all layer sets in the drawing, traverse the layer set based on preset layer name matching rules to filter and obtain a target layer, identify a table object from the target layer, and record the number of tables and the global coordinate range of the table for the table object, wherein the table object contains material information;

[0012] A text and cell extraction module is used to extract text elements and their spatial coordinate information in each identified table object through a first script to generate a text data file; and to extract cell boundary coordinates through a second script to generate a cell coordinate file;

[0013] a row and column index conversion module, configured to pair the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and convert the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file;

[0014] A row and column index and text mapping module is used to extract text data and cell data from the paired file, wherein the text data includes each text element and its spatial coordinates, and the cell data includes the cell boundary coordinates and its corresponding row and column indexes, and establish an association between text and cells based on the coordinates of the text elements and the cell boundary coordinates to form a mapping between the cell row and column indexes and the text;

[0015] The cell text splicing module is used to double-sort the text coordinates based on the mapping between the cell row and column indexes and the text, and to splice the sorted text to obtain the complete text content of the corresponding cell.

[0016] On the other hand, the present application also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned material information identification statistical method based on CAD export.

[0017] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the above-mentioned material information identification and statistical method based on CAD export.

[0018] In another aspect of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned statistical method for identifying material information based on CAD export.

[0019] The present invention parses CAD drawings and filters target layers based on rules to achieve automatic recognition and positioning of material tables. It combines dual scripts to extract text and cell boundaries and establish row and column index mapping to ensure accurate placement of material information. It introduces a deep table structure recognition model to automatically detect and complete missing boundary lines, effectively addressing problems such as incomplete table structure and drawing errors. Finally, it generates a formatted Excel file to achieve efficient, accurate, and structured output of material information, significantly improving statistical efficiency and reducing the error rate of manual operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0021] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0022] Figure 1 A flow chart of a material information identification and statistical method based on CAD export provided in an embodiment of the present invention.

[0023] Figure 2 This is a flowchart of row and column index conversion provided by an embodiment of the present invention.

[0024] Figure 3 A schematic diagram of the deep table structure recognition network model architecture provided by an embodiment of the present invention.

[0025] Figure 4 It is a structural schematic diagram of a material information identification and statistics device based on CAD export provided by an embodiment of the present invention.

[0026] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] The present application proposes a statistical method for identifying material information based on CAD export. The technical solution of the present application is described in detail below in conjunction with various embodiments.

[0029] like Figure 1 As shown, the embodiment of the present invention discloses a material information identification and statistics method 100 based on CAD export, including the following method steps:

[0030] Step S101: parse the DWG format engineering drawing to obtain all layer sets in the drawing, traverse the layer set based on the preset layer name matching rules to filter and obtain the target layer, identify the table object from the target layer, and record the number of tables and the global coordinate range of the table for the table object, wherein the table is a table containing material information.

[0031] In some embodiments, a DWG file is loaded and parsed through an API provided by the CAD software (e.g., AutoCAD's ActiveX interface), establishing an interactive channel with the drawing. The interface then extracts a collection of all layers in the drawing, including layer attributes such as name, visibility, and color. For example, an architectural drawing might contain multiple layers, such as "Outline," "Dimension," "Standard Drawing," and "Title Block," each carrying different types of metadata.

[0032] Next, the layer collection is traversed for name comparison based on pre-set layer name matching rules. For example, the specific logic is to obtain the name of each layer. If the name exactly matches the preset target layer name (for example, "Standard Drawing," where the standard drawing contains a target table, such as a material table) or meets fuzzy matching rules (for example, containing the keyword "material table"), the layer is marked as the target layer. After the screening is complete, only the target layer is retained for subsequent processing, and irrelevant layers (such as "Axis" and "Annotation" layers) are excluded to reduce data noise.

[0033] Next, from the target layer selected, all entities contained in the target layer are traversed, and the table object is identified by the entity type attribute (EntityType). In CAD, a table object usually corresponds to a specific type identifier (for example, AcDbTable), so the table object can be filtered by judging EntityType == "AcDbTable".

[0034] Finally, the total number of identified table objects is counted, and for each table object, the global coordinate range is extracted from its geometry attribute. For example, the left lower corner and the right upper corner coordinates of the table boundary are obtained (usually through the GetBoundingBox method), which are used to locate the spatial position of the table in the drawing. For example, the global coordinate range of a table is recorded as bottomLeft=(100, 200, 0), topRight=(500, 400, 0), indicating the coverage area of the table in the two-dimensional plane.

[0035] In step S102, for each identified table object, the first script extracts the text elements and their spatial coordinate information in the table to generate a text data file; and the second script extracts the cell boundary coordinates to generate a cell coordinate file.

[0036] In some embodiments, after identifying the table object in step S101, data extraction is performed for each table, and two special scripts are used to extract text content and cell boundary information, respectively. Based on the number of tables recorded in step S101 and the global coordinate range, each table object is processed in order according to the table ID (such as the unique handle Handle of the CAD entity). An independent processing context is created for each table object, including the table ID, global coordinate range, and other metadata, which are used for subsequent file naming and association. The script execution environment is initialized, the basic library required for script running is loaded through the LISP interpreter of the CAD software (such as the vl-load-com function of AutoCAD), and the script can call the entity operation interface of CAD.

[0037] Specifically, the first script is loaded, and the unique identifier [TableID] of the current table is passed in. The script locates the table object by TableID, and traverses all text entities in the table (including AcDbText single-line text and AcDbMText multi-line text). For each text entity, key information is extracted. For example, the text content Content: the text content (such as "bolt" "M10x50") is obtained through the TextString attribute; the text spatial coordinates X / Y / Z: the three-dimensional coordinates of the text insertion point (i.e., the spatial position of the text) are obtained through the InsertionPoint attribute.

[0038] For example, the extracted information is organized in JSON array format, with each record containing the "Content," "X," "Y," and "Z" fields. Files are saved using the naming convention [Filename]_[TableID]_TableText.json (for example, "Engineering Drawing A_12345_TableText.json"), and an incremental write mechanism is used (writing every 10 extracted text data items) to avoid memory overflows.

[0039] In this example, the second script is loaded and passed the current table's [TableID]. The script locates the table using the TableID, identifies the intersection of the table's horizontal and vertical lines, and determines the original bounds of each cell. The normalized bounds coordinates of each cell are recorded, illustratively including bottomLeft (the x and y values ​​of the lower left corner) and topRight (the x and y values ​​of the upper right corner).

[0040] Similarly, the extracted boundary information is organized in JSON format. Each record contains "bottomLeft" and "topRight" fields, each with "x" and "y" coordinate values. Files are saved using the naming convention [Filename]_[TableID]_RepairedTableCells.json (for example, "Engineering Drawing A_12345_RepairedTableCells.json"). Incremental writes are also used to ensure data security.

[0041] Optionally, perform basic validation on the two generated JSON files. Specifically, check whether the files are empty and whether key fields (such as the "X" coordinate of text and the "bottomLeft" field of a cell) are complete. If the files are corrupted or fields are missing, log an error (e.g., "The text file with table ID = 67890 is missing the Y coordinate field") and trigger the script to re-execute.

[0042] Step S103 : Pairing the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and converting the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file.

[0043] In some embodiments, in this step, the association basis between text and cells is established through file pairing, and the discrete boundary coordinates are converted into ordered row and column indexes to provide a structured reference for subsequent content integration.

[0044] Specifically, scan all JSON files in the storage directory and extract the file name prefix (the original DWG file name) and table ID for each file (e.g., from "Engineering Drawing A_12345_TableText.json," we parse the prefix "Engineering Drawing A" and the table ID "12345"). Text data files (TableText.json) and cell coordinate files (RepairedTableCells.json) with identical prefixes and table IDs are grouped together. For example, "Engineering Drawing A_12345_TableText.json" and "Engineering Drawing A_12345_RepairedTableCells.json" form a paired file.

[0045] Optionally, record the pairing relationship in dictionary form, with the key being the table ID and the value being the paths of the two files (for example, {"12345": ["path / text.json", "path / cell.json"]}). This ensures that subsequent steps can quickly call the corresponding file by table ID.

[0046] In some embodiments, converting the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file may, for example, specifically include:

[0047] Extract the x and y coordinates of the lower left corner and upper right corner of all cells from the cell coordinate file, and generate the XCoords set and the YCoords set. The XCoords set contains the x value of the lower left corner and the x value of the upper right corner of all cells, and the YCoords set contains the y value of the lower left corner and the y value of the upper right corner of all cells.

[0048] Sort the XCoords set in ascending order and remove duplicates. Calculate the difference between adjacent coordinates. When the difference is ≤ 0.5 coordinate units, merge them into the same coordinate to form the column boundary set ColBounds.

[0049] Sort the YCoords set in descending order and remove duplicates. Calculate the difference between adjacent coordinates. When the difference is ≤ 0.5 coordinate units, merge them into the same coordinate to form the row boundary set RowBounds.

[0050] Establish a mapping from coordinates to row and column indices. For the x coordinate, find the two adjacent column boundary intervals in ColBounds where it falls, and the corresponding column index is the index of the previous boundary of the interval in ColBounds; for the y coordinate, find the two adjacent row boundary intervals in RowBounds where it falls, and the corresponding row index is the index of the previous boundary of the interval in RowBounds.

[0051] Specifically, as shown in Figure 2 Step S201, the cell boundary coordinate set is extracted, and exemplarily, from the loaded cell coordinate file (RepairedTableCells.json), all cell objects are traversed to extract the x and y coordinates of the bottomLeft (lower left corner) and topRight (upper right corner) of each cell, and two basic coordinate sets are generated:

[0052] XCoords set: contains the x values of the lower left corners and the x values of the upper right corners of all cells, i.e. XCoords = {cell['bottomLeft']['x'], cell['topRight']['x'] for cell in cell list}. This set reflects all possible column boundary horizontal coordinates in the table.

[0053] YCoords set: contains the y values of the lower left corners and the y values of the upper right corners of all cells, i.e. YCoords = {cell['bottomLeft']['y'], cell['topRight']['y'] for cell in cell list}. This set reflects all possible row boundary vertical coordinates in the table.

[0054] The purpose of this step is to collect all potential boundary coordinate points in the table, providing original data for subsequent row and column boundary division.

[0055] Next, step S202, the generation of the column boundary set (ColBounds), specifically, the XCoords set is processed to form a standardized column boundary, which exemplarily includes:

[0056] First, arrange XCoords in ascending order (since the x-axis in the CAD coordinate system increases from left to right, ascending order conforms to the logic of columns from left to right), and remove duplicate values to obtain a preliminary ordered coordinate sequence;

[0057] Calculate the difference between adjacent coordinates after sorting. When the difference is ≤0.5 coordinate units, it is determined to be the boundary error of the same column (such as the slight offset when drawing in CAD), and it is merged into the same coordinate (usually taking the smaller value or the average value). For example, if the adjacent coordinates are 100.2 and 100.5, the difference is 0.3 (≤0.5), and they are merged into 100.2 (or 100.35);

[0058] The merged coordinate sequence is the column boundary set ColBounds, where each coordinate represents the left and right boundaries of a column in the table (e.g., ColBounds = [50.0, 100.2, 150.5] indicates that the table has 3 columns with column boundaries at 50.0, 100.2, and 150.5);

[0059] Next, step S203, generation of the row boundary set (RowBounds), which involves similar processing of the YCoords set to form normalized row boundaries, which exemplarily include:

[0060] Sort YCoords in descending order (since the y-axis in the CAD coordinate system increases from bottom to top, and descending order aligns with the logic of rows from top to bottom, e.g., the first row of the table is at the top with the largest y value), and remove duplicate values to obtain a preliminary ordered coordinate sequence.

[0061] Using the same merging logic as for column boundaries, calculate the difference between adjacent coordinates, and merge them into the same coordinate when the difference is ≤0.5 coordinate units. For example, adjacent coordinates 300.8 and 300.4 (difference 0.4) are merged into 300.8 (or 300.6).

[0062] The merged coordinate sequence is the row boundary set RowBounds, where each coordinate represents the top and bottom boundaries of a row in the table (e.g., RowBounds = [300.8, 250.3, 200.1] indicates that the table has 3 rows with row boundaries at 300.8, 250.3, and 200.1).

[0063] Next, step S204, establish a mapping relationship between coordinates and row-column indices, based on ColBounds and RowBounds, to convert any coordinate to the corresponding row-column index, which exemplarily includes:

[0064] Column index (ColIndex) calculation, specifically, for a certain x-coordinate (such as the X-coordinate of a text or the x-value of a cell boundary), determine its column by ColIndex = index(ColBounds, x). The specific logic is to find the two adjacent column boundary intervals that the x-coordinate falls into (e.g., ColBounds[i] ≤ x ≤ ColBounds[i+1]), and the column index corresponding to the x is i.

[0065] The row index (RowIndex) calculation is as follows: for a given y coordinate, determine its row using RowIndex = index(RowBounds, y). The specific logic is to find the interval between two adjacent row boundaries where the y coordinate falls (e.g., RowBounds[j] ≥ y ≥ RowBounds[j+1], since RowBounds is in descending order). The row index corresponding to y is j.

[0066] For example, if ColBounds = [50.0, 100.2, 150.5], and the X coordinate of a text is 70.0, which falls between 50.0 and 100.2, then its ColIndex = 0 (the first column); if RowBounds = [300.8, 250.3, 200.1], and the Y coordinate of a text is 280.0, which falls between 300.8 and 250.3, then its RowIndex = 0 (the first row).

[0067] Therefore, by normalizing the original coordinates, the boundary offset caused by CAD drawing errors is eliminated, and the physical coordinates are converted into logical row and column indexes, providing a unified reference standard for subsequent spatial matching of text content and cells.

[0068] Step S104: extract text data and cell data from the paired file, respectively, wherein the text data includes each text element and its spatial coordinates, and the cell data includes the cell boundary coordinates and its corresponding row and column indexes. An association between the text and the cell is established based on the coordinates of the text elements and the boundary coordinates of the cell, forming a mapping between the cell row and column indexes and the text.

[0069] In some embodiments, text data and cell data are extracted separately from the paired loaded JSON file. For example, the text data file ([file name]_[TableID]_TableText.json) is read and parsed to obtain detailed information of each text element, which for example includes: Content: text content (such as "bolt" "M10"); X / Y: spatial coordinates of the text in the CAD coordinate system (for positioning).

[0070] Next, read the cell coordinate file ([filename]_[TableID]_RepairedTableCells.json), and parse it to obtain the structured data of each cell in combination with the row and column indexes generated in step S103: bottomLeft / topRight: the coordinates of the lower left and upper right corners of the cell (x, y values); RowIndex / ColIndex: the row index and column index corresponding to the cell (such as (0,1) represents the first row and second column).

[0071] It is understandable that the text in the table is usually located inside or at the edge of the cell to which it belongs in the CAD drawing. Therefore, by determining whether the text coordinates fall within the boundary range of the cell, the association between the two can be achieved.

[0072] Next, determine whether the text belongs to the target cell. The exemplary implementation steps are as follows:

[0073] For each text element, extract its X / Y coordinates; for each cell, extract its bottomLeft.x / bottomLeft.y (left and bottom boundaries) and topRight.x / topRight.y (right and top boundaries).

[0074] Verify the coordinate range in both the horizontal and vertical directions. For example, the horizontal verification includes checking whether the text's X coordinate is within the cell's left and right boundaries (including tolerance):

[0075] cell['bottomLeft']['x'] - ε ≤ text['X'] ≤ cell['topRight']['x'] + ε

[0076] Wherein ε is the tolerance coefficient (default is 0.1 coordinate unit), which is compatible with the error of text slightly exceeding the boundary during CAD drawing (such as the text edge overlapping the cell line). It can be understood that the tolerance coefficient can be specifically set according to actual conditions and is not limited by the present invention.

[0077] The vertical verification includes checking whether the Y coordinate of the text is within the upper and lower bounds of the cell (including tolerance):

[0078] cell['bottomLeft']['y'] - ε ≤ text['Y'] ≤ cell['topRight']['y'] + ε

[0079] When the conditions are met in both the horizontal and vertical directions, the text is determined to belong to the cell, and the correspondence between the text and the cell is recorded.

[0080] Finally, the association results are converted into a structured mapping. For example, the mapping structure is defined as a dictionary with (RowIndex, ColIndex) as the key and a list of text elements as the value (e.g., {(0,1): [text1,text2], (1,2): [text3]}). For example, the mapping is temporarily stored in JSON format, containing the text content and original coordinates corresponding to each key.

[0081] Optionally, when a word matches multiple cells or a cell matches multiple words, regularization is used to ensure the accuracy of the association. For example,

[0082] Text matching multiple cells: Prioritize the cell with the smallest area (calculated by boundary coordinates: (topRight.x - bottomLeft.x) × (topRight.y - bottomLeft.y)), because text is usually located in the smallest enclosing cell.

[0083] Cell matching multiple words: temporarily store all matching words (subsequent step S105 merges them into complete content through sorting), without filtering, to avoid missing split text segments (such as "M10×50" is split into "M10" and "×50").

[0084] Step S105 , based on the mapping between the cell row and column indexes and the text, double sorting the text coordinates, and concatenating the sorted text to obtain the complete text content of the corresponding cell.

[0085] In some embodiments, based on the (RowIndex, ColIndex) → [text element list] mapping relationship generated in step S104, the cells are grouped and processed according to their row and column indexes:

[0086] Traverse each key-value pair of the mapping dictionary, where the key is (RowIndex, ColIndex) (such as (2,3) means the 3rd row and 4th column), and the value is a list of all text elements in the cell.

[0087] Extract the core information required for sorting from each text element, including Content (text content), X (horizontal coordinate), and Y (vertical coordinate).

[0088] For multiple paragraphs of text in the same cell (such as "M10" and "×50"), restore the natural reading order of the text by double sorting "vertically first, then horizontally":

[0089] Step 1: Sort by Y coordinate in descending order

[0090] Sort by the Y coordinate of the text from largest to smallest (sorted(key=lambda x: -x['Y'])). Since the larger the Y value in the CAD coordinate system, the higher the position, this sorting ensures the vertical order of the text from top to bottom. For example:

[0091] There are two lines of text in the cell, the first line Y=280 (content "bolt"), the second line Y=270 (content "GB / T5782"). After sorting, keep "bolt" in front and "GB / T 5782" in the back.

[0092] Step 2: Sort by X coordinate in ascending order

[0093] When the characters have the same Y coordinate (i.e., they are on the same horizontal line), they are sorted from smallest to largest by X coordinate (sorted(key=lambda x: x['X'])) to ensure the horizontal order of the characters from left to right. For example:

[0094] The text "M10" (X=120) and "×50" (X=140) in the same row are sorted and concatenated into "M10×50".

[0095] After sorting is completed, the complete contents of the cell are spliced ​​together according to the following rules, for example:

[0096] For the sorted text list, concatenate the contents sequentially. Horizontally concatenate text (e.g., "M10" + "×50" → "M10×50"), and vertically separate text with line breaks (\n) (e.g., "bolt" + "GB / T 5782" → "bolt\nGB / T 5782").

[0097] Optionally, special character processing includes removing redundant spaces (such as separating spaces automatically added by CAD), but retaining necessary format spaces (such as "5" → "5", "Specification: M10" → "Specification: M10").

[0098] Finally, the merged result is stored. For example, a final mapping dictionary of (RowIndex, ColIndex) → complete content is generated, such as {(2,3): "M10×50", (3,3): "Bolt\nGB / T 5782"}.

[0099] Therefore, through this step, the text in the CAD material table that was split due to length restrictions or typesetting requirements is reintegrated into readable complete content, while maintaining the original typesetting logic of the text (line breaks, order).

[0100] In some embodiments, after generating the complete cell content in step S105, the structured data is converted into a standardized Excel file, and the visualization and usability of the material information are achieved through workbook construction, data writing and format optimization.

[0101] Preferably, by establishing a workbook object, an independent worksheet is generated for each table in the DWG format engineering drawing; the complete text content of the corresponding cell is written into the corresponding cell in the independent worksheet according to the row and column index of the corresponding cell, and the cell is formatted, wherein the formatting includes cell style, column width, and header row formatting.

[0102] Specifically, for example, the Workbook() method of the OpenPyXL library is called to create a blank workbook object as a container for all table data. The default blank worksheet (Sheet) will be overwritten by subsequent operations to avoid redundancy. All tables identified in the DWG drawing (based on the number of tables and IDs recorded in step S101) are traversed, and a separate worksheet (Worksheet) is created for each table. Worksheet naming follows the "original file name prefix + table ID" rule, such as "Engineering Drawing A_12345". At the same time, the name length is limited to ≤31 characters (the upper limit of Excel worksheet names). The excess is replaced with an MD5 hash value (such as "Extended file name prefix..._8f3d7e") to ensure name uniqueness and compliance.

[0103] Next, the (RowIndex, ColIndex) → complete content mapping dictionary generated in step S105 is loaded, and the Excel cells are located sequentially according to the row and column indexes. For example, writing the complete content of the cell specifically includes:

[0104] Excel row and column indexes start at 1, so you need to increase the RowIndex and ColIndex in the mapping by 1 respectively (for example, RowIndex = 0 corresponds to Excel row number 1, and ColIndex = 1 corresponds to Excel column number 2).

[0105] Then, the complete content is written to the corresponding cell. For example, the content "Specification" of the cell (0,1) will be written to the cell A2 in Excel.

[0106] For content containing line breaks (such as "Bolt\nGB / T 5782"), write directly to preserve the original layout, and then use formatting settings to achieve automatic line breaks.

[0107] Preferably, the readability of Excel tables is improved through unified style settings, including:

[0108] Basic cell style settings include enabling word wrapping to ensure that content containing \n (such as multi-line text) is fully displayed; and setting vertical top alignment to conform to common table text layout conventions. This style is applied to all cells by default to ensure consistent formatting.

[0109] Automatically adjust column widths. Specifically, the system iterates through all cells in each column and calculates the character length of the content (Chinese characters are counted as 2 characters, English characters / numbers are counted as 1 character). Column widths are set based on the maximum character length within the column, within a range of 10-50 characters. If the maximum length is less than 10, the width is set to 10 (to avoid being too narrow); if the maximum length is greater than 50, the width is set to 50 (to avoid being too wide). In other cases, the width is set based on the actual maximum length (e.g., if the maximum length is 15, the column width is set to 15).

[0110] To enhance the style of the header row, specifically, locate the header row (RowIndex=0, corresponding to Excel row number 1) and apply a special style to all cells in this row:

[0111] The font is bold to highlight the distinction between the header and the content;

[0112] Fill the background color and set a light gray background to improve visual recognition.

[0113] Preferably, after writing and formatting all tables, save the workbook object as a local file named [Task ID]_output.xlsx ([Task ID] is a unique identifier assigned by the system for easy traceability). Temporarily save the file to a local temporary directory (such as / tmp / ) to prepare for subsequent upload to cloud storage (such as OSS) to avoid consuming memory resources.

[0114] In some embodiments, some material tables in drawings may contain missing boundary lines, which may cause confusion in boundary detection and coordinate mapping. To address this issue, this embodiment introduces a deep table structure recognition network to identify areas with missing boundary lines and complete the missing boundary lines based on the cell boundary coordinates extracted in step S102.

[0115] Preferably, based on the global coordinate range of the table, the corresponding table area image is obtained, the table area image and the cell boundary coordinates are input into a pre-trained deep table structure recognition network model, and a missing probability map of the table boundary line is output; based on the probability map and the cell boundary coordinates, the missing boundary lines are identified and completed, and the completed boundary coordinates are added to the cell coordinate file.

[0116] like Figure 3As shown, the deep table structure recognition network model includes an input layer, a feature extraction layer, a feature fusion layer, and an output layer. The input layer receives the table region image and the cell boundary coordinates, and outputs a preprocessed image. The preprocessing includes establishing a mapping relationship between the pixel coordinates of the table region image and the actual coordinates of the cell boundary, and removing noise points of the table region image. The feature extraction layer is a CNN layer, which receives the preprocessed image and outputs a feature map after convolution operation. The feature fusion layer is a feature pyramid network, which receives the feature map and outputs a fused second feature map. The output layer receives the second feature map, and outputs a missing probability map of the table boundary line through convolution operation and activation function.

[0117] Specifically, the specific network structure and implementation process include:

[0118] 1. Input layer and preprocessing

[0119] Input: Table region image of target layer in CAD drawing (size of 512x512 pixels grayscale image), containing identified table outline and part of boundary line.

[0120] Preprocessing: Remove noise points (such as residual drawing traces) in CAD drawing by Gaussian filtering, and retain edge features of table lines and cells; establish mapping between image pixel coordinates and cell boundary coordinates (such as bottomLeft and topRight) extracted in S102, to ensure that the network output is aligned with the actual coordinate system.

[0121] 2. Feature extraction network

[0122] Exemplarily, a lightweight convolutional neural network (CNN) is used to extract table edge features, which balances detection accuracy and efficiency. The network includes:

[0123] Convolution layer: Exemplarily, the first to third layers are 3x3 convolution kernels (channel numbers are 16→32→64 in turn), with a step of 1, and are combined with ReLU activation function to capture local edges of table lines (such as line segment endpoints and intersection points);

[0124] The fourth to fifth layers are 5x5 convolution kernels (channel numbers 128→256) with a step of 1, which capture global structural features of the table (such as row and column distribution rules and continuity of boundary lines);

[0125] Pooling operation: After every 2 layers of convolution, 2x2 max pooling (step 2) is performed to reduce the size of the feature map while retaining key features.

[0126] Output: Exemplarily, a feature map of 64x64x256, containing position, direction and integrity information of the table line.

[0127] 3. Multi-scale feature fusion layer

[0128] For example, the features of different levels are fused through the Feature Pyramid Network (FPN) to improve the positioning accuracy of missing boundary lines. Specifically, top-down fusion is performed: the global features of the high-level (64×64) are upsampled (bilinear interpolation) and spliced ​​with the local features of the middle-level (128×128) and low-level (256×256) layers to supplement the detailed information.

[0129] 4. Output Layer

[0130] This function is used to output missing probability maps for table boundary lines and locate areas requiring completion. For example, it includes two layers of 1×1 convolution (channel count 128→2), predicting the missing probabilities of row and column boundaries, respectively. Using a sigmoid activation function, it outputs two 512×512 probability maps (row missing probability map and column missing probability map) with pixel values ​​ranging from [0, 1] (higher values ​​indicate a greater likelihood of missing boundary lines at that location). The threshold is set to identify a missing boundary area when the pixel value is ≥ 0.7. It is understood that the threshold can be set based on actual conditions and is not a limitation of this invention.

[0131] In some embodiments, identifying and completing missing boundary lines includes extracting pixel coordinates of row / column missing areas from the missing probability map, converting them into actual coordinates based on the mapping relationship, and automatically completing the row / column missing areas based on the average spacing between adjacent rows / columns of the cell boundary coordinates.

[0132] Specifically, the pixel coordinates of the missing row / column areas are extracted from the probability map output by the network and converted into actual CAD coordinates (based on the coordinate mapping relationship of the input layer);

[0133] Among them, the categorized missing types include, for example, row missing, specifically, the upper / lower boundary coordinates of multiple consecutive cells are discontinuous (for example, the difference between topRight.y of a row and bottomLeft.y of the next row far exceeds the normal row spacing); for example, column missing, specifically, the left / right boundary coordinates of multiple consecutive cells are discontinuous (for example, the difference between topRight.x of a column and bottomLeft.x of the next column is abnormal).

[0134] From the cell boundary coordinates extracted in step S102, select complete and continuous boundaries as the completion benchmark:

[0135] Sort the YCoords set (the y coordinates of the cells) in descending order, calculate the difference between adjacent coordinates (normal row spacing), and remove outliers that deviate from the mean by more than 2 times;

[0136] Sort the XCoords collection (cell x-coordinates) in ascending order, calculate the difference between adjacent coordinates (normal column spacing), and retain the coordinate sequences of more than 3 consecutive coordinates that meet the mean.

[0137] In this embodiment, for the missing row area, the average spacing between adjacent valid rows is taken (e.g., the average row height of the top three rows is 10 coordinate units). Based on the starting y coordinate of the missing area, the missing row boundaries are sequentially calculated according to the average (e.g., starting at y=200, the boundaries at y=190 and 180 are filled in).

[0138] For the missing column area, take the average of the spacing between adjacent valid columns (for example, the average column width of the two columns on the left is 15 coordinate units); based on the starting x coordinate of the missing area, calculate the missing column boundaries according to the average (for example, starting at x=100, fill in the boundaries of x=115 and 130).

[0139] Finally, the completed boundary coordinates are added to the XCoords / YCoords set, and the sorting, deduplication and merging operations are re-performed to generate complete ColBounds and RowBounds. For details, please refer to the above embodiment and will not be repeated here.

[0140] Figure 4 A material information identification and statistics device 400 based on CAD export is shown. Figure 1 Corresponding to the method embodiment shown, the device can be applied to various electronic devices. Specifically, it includes:

[0141] The table parsing and identification module 401 is configured to parse a DWG-formatted engineering drawing to obtain a set of all layers in the drawing, traverse the set of layers based on a preset layer name matching rule to filter out a target layer, identify a table object from the target layer, and record the number of tables and the global coordinate range of the table for each table object, wherein the table object contains material information.

[0142] The text and cell extraction module 402 is configured to extract text elements and their spatial coordinate information in each identified table object using a first script to generate a text data file; and to extract cell boundary coordinates using a second script to generate a cell coordinate file;

[0143] a row and column index conversion module 403 for pairing the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and converting the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file;

[0144] The row and column index and text mapping module 404 is configured to extract text data and cell data from the paired file, wherein the text data includes each text element and its spatial coordinates, and the cell data includes the cell boundary coordinates and their corresponding row and column indexes, establish an association between the text and the cell based on the coordinates of the text elements and the cell boundary coordinates, and form a mapping between the cell row and column index and the text;

[0145] The cell text splicing module 405 is used to perform a double sort on the text coordinates based on the mapping between the cell row and column indexes and the text, and splice the sorted text to obtain the complete text content of the corresponding cell.

[0146] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0147] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes, but is not limited to, various terminal devices such as computers, mobile phones, tablets, smart watches, and wristbands. The network device includes, but is not limited to, network hosts, single network servers, multiple network server sets, or a collection of cloud computing-based computers, and can be used to implement some of the processing functions required for setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computers.

[0148] Figure 5 The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 into the random access memory (RAM) 503. Various programs and data required for system operation are also stored in the RAM 503. The CPU 501, ROM 502 and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0149] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker; a storage section 508 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet.

[0150] In particular, the methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the method of the present application are performed.

[0151] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.

[0152] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by an instruction execution system, device or device or used in combination with it.

[0153] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0154] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0155] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0156] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

Claims

1. A statistical method for identifying material information based on CAD export, characterized in that: The following steps are involved: Step S101: Parse a DWG format engineering drawing to obtain all layer sets in the drawing, traverse the layer set based on a preset layer name matching rule to filter and obtain a target layer, identify a table object from the target layer, and record the number of tables and the global coordinate range of the table for the table object, wherein the table object contains material information; Step S102: for each identified table object, extract the text elements and their spatial coordinate information in the table using a first script to generate a text data file; Extract the cell boundary coordinates through the second script and generate a cell coordinate file; Step S103, pairing the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and converting the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file; Step S104: extracting text data and cell data from the paired file, wherein the text data includes each text element and its spatial coordinates, and the cell data includes cell boundary coordinates and their corresponding row and column indexes. An association is established between the text and the cell based on the coordinates of the text elements and the cell boundary coordinates, forming a mapping between the cell row and column indexes and the text. Step S105 , based on the mapping between the cell row and column indexes and the text, double sorting the text coordinates, and concatenating the sorted text to obtain the complete text content of the corresponding cell.

2. The material information identification and statistical method based on CAD export according to claim 1, characterized in that: Also includes, By creating a workbook object, an independent worksheet is generated for each table in the DWG format engineering drawing; The complete text content of the corresponding cell is written into the corresponding cell in the independent worksheet according to the row and column index of the corresponding cell, and the cell is formatted, wherein the formatting includes cell style, column width, and header row formatting.

3. The material information identification and statistical method based on CAD export according to claim 1, characterized in that: Based on the global coordinate range of the table, a corresponding table area image is obtained, the table area image and the cell boundary coordinates are input into a pre-trained deep table structure recognition network model, and a missing probability map of the table boundary line is output; Based on the probability map and the cell boundary coordinates, missing boundary lines are identified and completed, and the completed boundary coordinates are added to the cell coordinate file.

4. The material information identification and statistical method based on CAD export according to claim 3 is characterized in that: The deep table structure recognition network model includes an input layer, a feature extraction layer, a feature fusion layer, and an output layer. The input layer receives the table area image and the cell boundary coordinates and outputs a preprocessed image, wherein the preprocessing includes establishing a mapping relationship between the pixel coordinates of the table area image and the actual coordinates of the cell boundary and removing noise from the table area image; The feature extraction layer is a CNN layer, which receives the preprocessed image and outputs a feature map after convolution operation; The feature fusion layer is a feature pyramid network, which receives the feature map and outputs a fused second feature map; The output layer receives the second feature map and outputs the missing probability map of the table boundary line through a convolution operation and an activation function.

5. The material information identification and statistical method based on CAD export according to claim 4 is characterized in that: The identifying and completing missing boundary lines includes extracting pixel coordinates of row / column missing areas from the missing probability map, converting them into actual coordinates based on the mapping relationship, and automatically completing the row / column missing areas based on the average of the spacing between adjacent rows / columns of the cell boundary coordinates.

6. The material information identification and statistical method based on CAD export according to claim 3 is characterized in that: The cell boundary coordinates in the table are converted into ordered row and column indexes based on the pairing file, including: Extracting the x and y coordinates of the lower left corner and the upper right corner of all cells from the cell coordinate file to generate an XCoords set and a YCoords set, wherein the XCoords set contains the x values ​​of the lower left corner and the x values ​​of the upper right corner of all cells, and the YCoords set contains the y values ​​of the lower left corner and the y values ​​of the upper right corner of all cells; Sort the XCoords set in ascending order and remove duplicates. Calculate the difference between adjacent coordinates. When the difference is ≤ 0.5 coordinate units, merge them into the same coordinate to form the column boundary set ColBounds. Sort the YCoords set in descending order and remove duplicates. Calculate the difference between adjacent coordinates. When the difference is ≤ 0.5 coordinate units, merge them into the same coordinate to form the row boundary set RowBounds. Establish a mapping from coordinates to row and column indices. For the x coordinate, find the two adjacent column boundary intervals in ColBounds where it falls, and the corresponding column index is the index of the previous boundary of the interval in ColBounds; for the y coordinate, find the two adjacent row boundary intervals in RowBounds where it falls, and the corresponding row index is the index of the previous boundary of the interval in RowBounds.

7. A material information identification and statistics device based on CAD export, characterized in that: include: A table parsing and identification module is used to parse a DWG format engineering drawing to obtain all layer sets in the drawing, traverse the layer set based on preset layer name matching rules to filter and obtain a target layer, identify a table object from the target layer, and record the number of tables and the global coordinate range of the table for the table object, wherein the table object contains material information; A text and cell extraction module, configured to extract text elements and their spatial coordinate information in each identified table object through a first script, and generate a text data file; Extract the cell boundary coordinates through the second script and generate a cell coordinate file; a row and column index conversion module, configured to pair the text data file and the cell coordinate file based on the file name prefix and the table ID to obtain a pairing file, and convert the cell boundary coordinates in the table into ordered row and column indexes based on the pairing file; A row and column index and text mapping module is used to extract text data and cell data from the paired file, wherein the text data includes each text element and its spatial coordinates, and the cell data includes the cell boundary coordinates and its corresponding row and column indexes, and establish an association between text and cells based on the coordinates of the text elements and the cell boundary coordinates to form a mapping between the cell row and column indexes and the text; The cell text splicing module is used to double-sort the text coordinates based on the mapping between the cell row and column indexes and the text, and to splice the sorted text to obtain the complete text content of the corresponding cell.

8. An electronic device, wherein: include: at least one processor; and a memory in communication with the processor; wherein, The memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable medium having computer program instructions stored thereon, characterized in that: The computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Spreadsheet structured recognition and extraction method based on CAD basic elements

    CN112241411A

  • Method and device for converting table in image into spreadsheet

    CN113688795A

  • Coordinate labeling and coordinate exporting method and system of terminal equipment based on CAD tool, electronic equipment and storage medium

    CN118965477A

  • Method and system for adjusting CAD drawing segmentation parameters in combination with region recognition

    CN120374991A