An unstructured CAD table extraction method and system
Patent Information
- Application Number
- CN202611023113.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-10
AI Technical Summary
[0009]针对现有技术中的上述不足,本发明提供的一种非结构化CAD表格提取方法及系统解决了现有技术对于嵌套块参照与复合表格等异构封装格式的底层穿透解析能力匮乏、过度依赖人工预设的硬编码容差阈值,导致算法普适性极差、计算资源消耗大、缺乏有效的全局降噪机制与精准的拓扑结构还原能力的问题
1、本发明能够彻底摒弃对人工经验阈值的依赖,转而纯粹基于目标图纸底层的内生数据统计分布特征,进行全域空间算子的自适应动态推演;进而在克服物理绘制断线与坐标系畸变的前提下,实现由离散的二维物理坐标系向高保真逻辑网格拓扑的降维重构与精准映射,确保对非结构化CAD表格的准确、快速提取。
Smart Images

Figure CN122531053B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of CAD file content extraction, specifically to a method and system for extracting unstructured CAD tables. Background Technology
[0002] In fields such as engineering design, architectural drafting, and mechanical manufacturing, designers extensively use CAD software for drawing. These drawings typically contain numerous tables to represent key data such as parts lists, material lists, and window / door schedules. However, due to historical practices or software limitations, many tables in CAD drawings are not composed of native structured table entities (such as the Table object in AutoCAD), but rather are "unstructured tables" (or "scattered table tables") visually aligned and pieced together from discrete line segments, polylines, nested block references, and independent text fragments (Text / MText).
[0003] Exporting unstructured, highly heterogeneous CAD table data into structured spreadsheets (such as Excel) is a long-standing common need in the industry. Currently, existing extraction and conversion technologies mainly include the following solutions, but all have extremely obvious limitations: Existing technology 1: Topological method based on strict geometric closed region detection This type of method traverses all line segments in the drawing, searching for strictly closed polygons (such as rectangles) formed by connecting the beginning and end, and uses these as the physical cells of the table. Problems exist (lack of fault tolerance and noise reduction capabilities): It is extremely dependent on the drafter's adherence to standards. In actual engineering drawings, due to drawing errors, there are often tiny gaps between lines, endpoints intersecting, or a single long line segment running through multiple rows and columns. This strict topological closure algorithm has extremely low fault tolerance; even a 0.1mm drawing deviation will cause the entire cell to fail to be recognized. Furthermore, this method cannot handle minor coordinate fluctuations, easily producing misaligned "fragmented meshes" during export.
[0004] Existing technology 2: Hard-coded matching method based on fixed tolerance and absolute coordinates To address the aforementioned line breakage problem, some existing technologies have introduced a coordinate distance matching method. This involves setting a fixed empirical tolerance threshold (e.g., a fixed allowable break gap of 2.0 mm or a tilt angle of 1.0 degrees) and calculating the geometric inclusion relationship between entities within this fixed tolerance. However, this method has drawbacks (lack of adaptability and distortion resistance): Since the drawing scales of actual engineering drawings vary greatly and are often accompanied by unpredictable global scaling or local rotation, the method of manually pre-setting a fixed tolerance (hard-coded) presents insurmountable limitations. In drawings with enlarged scales, the fixed stitching tolerance may be too small, leading to stitching failure (missed detection); in drawings with reduced scales, the fixed tolerance may be too large, incorrectly merging adjacent dense rows and columns (false detection). This method completely lacks adaptability to different drawing noise levels; furthermore, when faced with coordinate system distortion caused by a slight overall deviation in the drawing, the bounding box algorithm based on absolutely orthogonal coordinates will completely fail due to mismatch.
[0005] Existing technology 3: Reconstruction method based on image rasterization and computer vision (CV) This type of method first rasterizes the native vector graphics of CAD drawings into an image format, and then uses object detection models and optical character recognition (OCR) technology from computer vision to extract table topology and text content. However, it has drawbacks (precision inverse dimensionality reduction and massive computational cost): This method completely abandons the native, absolutely precise vector coordinates and plain text string data of the CAD underlying layer, "inversely reducing" the high-dimensional, high-precision vector graphics into bitmap data with irreversible pixel loss. This process not only introduces extremely high system computational and memory overhead, but the OCR algorithm also has an inherent character misrecognition rate, especially when dealing with complex professional special symbols (such as rebar grade symbols) and superscripts and subscripts that frequently appear in engineering drawings, which can easily lead to fatal data tampering. Furthermore, for ultra-large engineering drawings containing tens of thousands of entities, the image recognition model often crashes due to memory overflow or misses key local details due to resolution compromises, failing to meet the large-scale, high-frequency parsing requirements of industrial applications.
[0006] For example, Chinese patent CN116933350A, entitled "Method, Apparatus, Device and Storage Medium for List Compilation Based on Drawing Tables," discloses a method for obtaining line primitives from CAD drawings, assembling them to generate table cells, and then matching text primitives with the cells. Although this method attempts vector extraction, its underlying "line assembly logic" is a typical rigid topology method, failing to address the mesh reconstruction problem when line segments have breakage tolerances and intersection burrs. When dealing with scattered line tables with significant drawing noise in practical engineering, the lack of an adaptive noise reduction mechanism easily leads to the generation of numerous fragmented meshes or assembly failures.
[0007] For example, Chinese patent CN111553187B, entitled "Method and System for Recognizing Tables in CAD Drawings," discloses a method for analyzing line features using an artificial intelligence model to identify table data. This type of AI or machine vision-based solution thoroughly exposes the shortcomings of "Primary Technology Three": not only does it consume significant computational resources and introduce unnecessary black-box problems in model inference, but it also struggles to adaptively address the issues of minor coordinate system distortion correction and topological absolute reconstruction of complex cross-column merged cells based on the inherent pure vector precision of the drawing.
[0008] In summary, existing CAD table parsing and extraction technologies generally suffer from the following insurmountable technical bottlenecks: First, it lacks the ability to penetrate and resolve heterogeneous encapsulation formats such as nested block references and composite tables at the lowest level; Second, the over-reliance on manually preset hard-coded tolerance thresholds results in extremely poor algorithm universality, making it unable to withstand the robustness risks caused by global scaling of drawings or local coordinate distortion. Third, it consumes a large amount of computing resources; Fourth, when faced with physical drawing errors, slight coordinate jitter, and complex cross-row and cross-column cell merging that are common in real engineering drawings, there is a lack of effective global noise reduction mechanisms and accurate topology restoration capabilities. Summary of the Invention
[0009] To address the aforementioned shortcomings in existing technologies, this invention provides an unstructured CAD table extraction method and system that solves the problems of insufficient low-level penetration and parsing capabilities for heterogeneous encapsulation formats such as nested block references and composite tables, excessive reliance on manually preset hard-coded tolerance thresholds, resulting in extremely poor algorithm universality, high computational resource consumption, and a lack of effective global noise reduction mechanisms and accurate topology restoration capabilities.
[0010] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for extracting unstructured CAD tables is provided, which includes the following steps: Receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and perform heterogeneous entity deconstruction and spatial feature vectorization, outputting a set of line segment entities and a set of text entities; For a set of text entities, the baseline visual proportion of the drawing is dynamically extracted based on the statistical frequency distribution principle, thereby obtaining the fluctuation range of font height in the current drawing; For the set of line segments, the global tolerance of physical breakage and trimming error is adaptively calculated based on the spatial nearest neighbor statistical algorithm, and then the angle feature filtering is performed to obtain the set of horizontal lines and the set of vertical lines. Using the geometric center of each text entity as the origin of the probe, orthogonal rays are initiated in the set of horizontal and vertical lines. Then, the boundaries of each text entity are captured by combining the font height fluctuation range, and a logically closed virtual bounding box is generated for the text entity. Projected coordinates are constructed from the vertical and horizontal boundary coordinates of all virtual bounding boxes. Then, the discrete virtual bounding box boundary coordinates are reconstructed into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner by one-dimensional streaming mean clustering. The physical boundary coordinates of the virtual bounding box are mapped to the orthogonal full-screen mathematical grid sequence for dimensionality reduction. Text entities with the same mapping relationship are grouped into the same logical cell, and the basic two-dimensional logical coordinates and logical span of the text entity are calculated. At the same time, the topological parameters of the corresponding logical cell are obtained. The physical column width and row height parameters of the target spreadsheet are generated based on the coordinate difference mapping of adjacent elements in the orthogonal full-screen mathematical grid sequence; the cells of the target spreadsheet are topologically merged based on the basic two-dimensional logical coordinates and logical span of the text entities; Text entities within the same logical cell are sorted and concatenated in a dimensionality-reduced manner. The topological parameters of the corresponding logical cells are then used to fill the logical coordinates of the target spreadsheet and serialized for output, thus completing the extraction of unstructured CAD tables.
[0011] A system based on an unstructured CAD table extraction method is provided, comprising: The first module is used to receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and to perform heterogeneous entity deconstruction and spatial feature vectorization, and output the line segment entity set and the text entity set. The second module is used to dynamically extract the baseline visual proportion of the drawing based on the statistical frequency distribution principle for the text entity set, and then obtain the font height fluctuation range in the current drawing; for the line segment entity set, it adaptively calculates the global tolerance of physical break and trimming error based on the spatial nearest neighbor statistical algorithm, and then performs angle feature filtering to obtain the horizontal line set and the vertical line set. The third module is used to initiate orthogonal rays in the set of horizontal and vertical lines with the geometric center of each text entity as the probe origin, and then capture the boundary of each text entity by combining the font height fluctuation range, and generate a logically closed virtual bounding box for the text entity. The fourth module is used to construct projected coordinates for the vertical and horizontal boundary coordinates of all virtual bounding boxes, and then reconstruct the discrete virtual bounding box boundary coordinates into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner through one-dimensional streaming mean clustering. The fifth module is used to perform dimensionality reduction mapping of the physical boundary coordinates of the virtual bounding box to the orthogonal full-screen mathematical grid sequence, group text entities with the same mapping relationship into the same logical cell, calculate the basic two-dimensional logical coordinates and logical span of the text entity, and obtain the topological parameters of the corresponding logical cell. The sixth module is used to generate the physical column width and row height parameters of the target spreadsheet based on the coordinate difference mapping of adjacent elements in the orthogonal full-screen mathematical grid sequence; and to perform topological merging of the cells of the target spreadsheet based on the basic two-dimensional logical coordinates and logical span of the text entity. The seventh module is used to sort and concatenate text entities within the same logical cell in a dimensionality reduction manner, fill them into the logical coordinates of the target spreadsheet with the topological parameters of the corresponding logical cell, and serialize and output them to complete the extraction of unstructured CAD tables.
[0012] The beneficial effects of this invention are as follows: 1. This invention can completely eliminate the reliance on human experience thresholds and instead rely purely on the endogenous data statistical distribution characteristics of the target drawing to perform adaptive dynamic deduction of global spatial operators; thereby, under the premise of overcoming physical drawing line breaks and coordinate system distortion, it realizes the dimensionality reduction reconstruction and accurate mapping from discrete two-dimensional physical coordinate system to high-fidelity logical grid topology, ensuring accurate and fast extraction of unstructured CAD tables.
[0013] 2. Addressing the problem of existing technologies relying excessively on manually preset hard-coded tolerance thresholds, resulting in poor algorithm universality, this invention abandons the traditional hard-coded tolerance of fixed values or arithmetic means. On one hand, through spatial nearest neighbor statistics and text height mode extraction, a dynamically derived stitching operator can adaptively scale with the drawing's proportions, successfully bridging gaps that are visually plausible but physically broken. On the other hand, combining angle filtering based on the normal distribution 3σ criterion and a one-dimensional streaming mean clustering algorithm, discrete physical lines with minute displacement errors and angular distortions are forcibly converged into an absolutely orthogonal mathematical logic grid. This mechanism completely eliminates interference from manual drawing errors and scale heterogeneity, significantly reducing the initial costs of drawing standardization.
[0014] 3. Addressing the shortcomings of existing technologies in handling heterogeneous encapsulation formats such as nested block references and composite tables, including insufficient low-level penetration resolution and a lack of effective global noise reduction mechanisms and accurate topology reconstruction capabilities, this invention uses the text's geometric center as the absolute origin and performs ray collision optimization within orthogonal space. Specifically, it introduces a dynamic depth constraint threshold based on the drawing's geometric noise floor. It can accurately penetrate invalid, free-floating interference primitives, lock onto the true structural sidewalls, and thus dynamically generate an absolutely closed "virtual bounding box." Combined with the subsequent "extreme approximation space mapping mechanism," it completely eliminates the dependence on the geometric appearance of the graphics, directly achieving this at the pure mathematical matrix level through logical index differences (…). and The system accurately inverts and derives the merged attributes of cells. This allows the exported target spreadsheet to achieve 100% high-fidelity, lossless restoration of the original CAD drawing in terms of row and column structure and span topology.
[0015] 4. To address the issue of high computational resource consumption in existing technologies, this invention implements a two-stage core algorithm dimensionality reduction strategy: First, in the spatial analysis pre-stage, statistical parallelism tolerance is used to pre-classify line segments into sets of orthogonal horizontal lines and vertical lines, instantly halving the single-step retrieval space for ray optimization; second, in the coordinate-to-logic grid mapping stage, global linear traversal is completely abandoned, and a heuristic extreme value approximation retrieval mechanism based on sequence monotonicity is adopted, utilizing the "bottoming out" characteristic of absolute differences to trigger early termination instructions. The deep coupling of these two core technologies significantly reduces ineffective computational branches. When processing ultra-large, dense engineering drawings, the analytical performance achieves an order-of-magnitude (tens of times) performance leap compared to traditional topology methods, while maintaining extremely low memory resource consumption.
[0016] 5. In the structured output stage, this invention innovatively introduces a "strict dimensionality reduction sorting mechanism based on two-dimensional physical coordinates" for multi-line discrete text fragments mapped to the same logical cell coordinates. By setting the primary sorting key as the Y-axis geometric center coordinate (decreasing from top to bottom) and the secondary sorting key as the X-axis geometric center coordinate (increasing from left to right), it forcibly reconstructs a one-dimensional linear sequence, completely shielding the interference of the underlying handle order and perfectly restoring the natural visual reading order of humans. Based on this, this invention further embeds a data purification and adaptive encapsulation engine: First, it performs a hash-matching-based filtering operation on the one-dimensional linear sequence, completely eliminating the overlapping redundancy caused by in-situ copying of primitives; second, it introduces a data type adaptive conversion mechanism to intelligently identify and implicitly convert numerical floating-point data; finally, by pre-initializing the full logical grid base map and performing strict memory occupancy, it fundamentally eliminates the problem of physical border loss caused by cross-row and cross-column topology merging. This series of deeply coupled end-point defense mechanisms opens up the process of delivering unstructured data to highly available structured data. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the method. Figure 2 This is a technical roadmap for this method; Figure 3This is a schematic diagram of an unstructured CAD table in the embodiment. Detailed Implementation
[0018] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0019] like Figure 1 and Figure 2 As shown, this method for extracting unstructured CAD tables includes the following steps: S1. Receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and perform heterogeneous entity deconstruction and spatial feature vectorization, outputting a set of line segment entities and a set of text entities. S2. For a set of text entities, the baseline visual proportion of the drawing is dynamically extracted based on the statistical frequency distribution principle, thereby obtaining the font height fluctuation range in the current drawing. S3. For the set of line segments, the global tolerance of physical line breakage and trimming error is adaptively calculated based on the spatial nearest neighbor statistical algorithm, and then the angle feature filtering is performed to obtain the set of horizontal lines and the set of vertical lines. S4. Using the geometric center of each text entity as the origin of the probe, initiate orthogonal rays in the set of horizontal lines and the set of vertical lines, and then combine the font height fluctuation range to capture the boundary of each text entity, generating a logically closed virtual bounding box for the text entity. S5. Construct projected coordinates for the vertical and horizontal boundary coordinates of all virtual bounding boxes respectively. Then, reconstruct the discrete virtual bounding box boundary coordinates into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner using one-dimensional streaming mean clustering. S6. Perform dimensionality reduction mapping on the physical boundary coordinates of the virtual bounding box to the orthogonal full-screen mathematical grid sequence, group text entities with the same mapping relationship into the same logical cell, calculate the basic two-dimensional logical coordinates and logical span of the text entity, and obtain the topological parameters of the corresponding logical cell. S7. Generate the physical column width and row height parameters of the target spreadsheet based on the coordinate difference mapping of adjacent elements in the orthogonal full-screen mathematical grid sequence; perform topological merging of the cells of the target spreadsheet based on the basic two-dimensional logical coordinates and logical span of the text entity; S8. Perform dimensionality reduction sorting and concatenation on the text entities within the same logical cell, fill them into the logical coordinates of the target spreadsheet using the topology parameters of the corresponding logical cells, and serialize and output them to complete the extraction of unstructured CAD tables.
[0020] In the specific implementation process, the methods for receiving the unstructured CAD table entity objects selected by the user in the target CAD drawing and performing heterogeneous entity deconstruction and spatial feature vectorization include: S1-1: Receive the unstructured CAD table entity object to be exported selected by the user in the target CAD drawing, and extract the underlying type identifier of the unstructured CAD table entity object; S1-2. For unstructured CAD table entity objects whose underlying type is a composite encapsulated entity, call the CAD underlying decomposition interface to perform physical decomposition operations on the composite encapsulated entity to generate a discrete set of sub-entities; among which the composite encapsulated entity includes native table, block reference or two-dimensional polyline (Polyline2d). S1-3. Recursively call the physical decomposition algorithm with each sub-entity in the sub-entity set as input parameter to obtain the basic primitive features contained in the composite encapsulated entity, and simultaneously release the memory resources occupied by the sub-entity; where the basic primitive features are lines or text. S1-4. If the basic graphic element feature is a straight line, then store it in a preset set of line segment entities, denoted as... ,in Represents a set of line segment entities. Represents the first line segment in the set of line segment entities. Each line segment From the starting coordinates and endpoint coordinates Composition; if the basic primitive feature is a text entity, then extract the height of each text entity. ,content Calculate the geometric center coordinates of the bounding box of the text entity. The text entities are then stored in a predefined text entity set, represented as... ,in A collection of text entities For the first text entity in the set A text entity; the geometric center coordinates of the bounding box of the text entity. That is, the geometric center coordinates of the text entity, which is the absolute origin for subsequent spatial exploration; S1-5. For unstructured CAD table entity objects whose underlying type is a standard polyline, traverse its vertex sequence and extract the sub-segment type formed by adjacent vertices; when the sub-segment type is a straight line segment, extract the spatial coordinate data of the current vertex and the next adjacent point, convert them into independent basic line segment features and store them in the preset line segment entity set.
[0021] In the specific implementation process, the specific methods for dynamically extracting the baseline visual proportion of the drawing based on the statistical frequency distribution principle, and then obtaining the fluctuation range of font height in the current drawing, include: S2-1. Extract the height of all text entities in the text entity set, and then construct a global text height sample sequence; S2-2. To address the inconsistency in absolute scale caused by differences in measurement units (such as millimeters and meters) between different target CAD drawings, this step calculates the initial average height of the global text height sample sequence and sets 5% to 10% of the initial average height as the height discretization step size. S2-3. According to the height discretization step size, map the global text height sample sequence to the corresponding height interval, count the entity frequency in each height interval, and generate a height frequency histogram. S2-4. Perform peak retrieval in the height frequency histogram to identify the highest peak interval. Based on the prior engineering rule that real CAD table text typically uses a uniform standard font height, extract the central tendency feature (such as the mode or median) of the samples within this highest peak interval and use it as the global baseline visual scale for the drawing. ; S2-5. Calculate the standard deviation of the sample proportions deviating from the global baseline visual proportions within the maximum peak interval. This range is used as the font height fluctuation range in the current drawing, thereby quantifying the font height fluctuation range caused by changes in hierarchical headings or manual scaling in the current drawing.
[0022] In some embodiments, the highly discretized step size can also be set to an absolute physical unit range of [0.1, 1.0] based on expert experience.
[0023] In the specific implementation process, the global tolerance of physical line breakage and trimming errors is adaptively calculated based on the spatial nearest neighbor statistical algorithm, and then angular feature filtering is performed to obtain the set of horizontal lines and the set of vertical lines. The specific methods include: S3-1, Calculate the set of line segment entities The inclination angles of each line segment are then used to construct horizontal and vertical angle density histograms, respectively. S3-2. Use the angle of the highest main peak in the horizontal angle density histogram as the reference horizontal angle. And calculate the standard deviation of the horizontal angles in the horizontal angle density histogram. ; S3-3. Use the angle of the highest main peak in the vertical angle density histogram as the reference vertical angle. And calculate the standard deviation of the vertical angle in the vertical angle density histogram. ; S3-4. Based on the Laida criterion, generate the horizontal parallelism tolerance. and vertical parallelism tolerance ;in To prevent the tolerance from reaching a minimum of zero, the horizontal parallelism tolerance and the vertical parallelism tolerance constitute the global tolerance. S3-5. Perform angular feature filtering on the line segment entity set using global tolerance to obtain the horizontal line set and the vertical line set. The corresponding expression is:
[0024]
[0025] in A set of horizontal lines; A set of vertical lines; representing; Represents a set of line segment entities The first in A single line segment entity; Line segment entity The angle of inclination.
[0026] In the specific implementation process, taking the geometric center of each text entity as the probe origin, orthogonal rays are initiated from the sets of horizontal and vertical lines. Then, the boundaries of each text entity are captured by combining the font height fluctuation range. The specific method for generating a logically closed virtual bounding box for the text entity includes: S4-1, Traversing the Set of Line Segments For each independent line segment in the drawing, calculate the Euclidean distance from the endpoint of each line segment to the endpoints of other non-collinear line segments in its spatial neighborhood; after removing completely closed points with zero Euclidean distance, construct a sample set of endpoint spacing that characterizes the scale of the physical break in the drawing. S4-2. Sort the endpoint spacing sample set in ascending order and extract the median in the middle of the sequence. Since the median has a strong statistical resistance to extreme free-floating lines (maximum values) or overlapping noise points (minimum values) in the drawing, this embodiment uses the median as the adaptive break spacing for the overall broken line defect scale of the current drawing. And based on the median at the very center of the sequence, the standard deviation of the sequence is obtained, that is, the standard deviation corresponding to the adaptive fracture spacing; S4-3. Weight the reference visual scale of the drawing with the adaptive break spacing to obtain the dynamic stitching operator. Its expression is:
[0027] in , These are the normalized feature fusion weight coefficients, and ; S4-4, Traversing the Collection of Text Entities Using the geometric center of each text entity as the probe origin, in the set of horizontal lines With vertical line set Initiate orthogonal rays to obtain the X-axis projection domain of individual horizontal line segments in the set of horizontal lines. Y-axis projection domain of a single vertical line segment in the set of vertical lines ;in and These are the minimum and maximum coordinate values of the X-axis projection domain of a single horizontal line segment, respectively. and These are the minimum and maximum coordinate values of the Y-axis projection domain for a single vertical line segment, respectively. S4-5, X-axis projection domain of a single horizontal line segment Based on this, a dynamic stitching operator is introduced. Construct extended collision range ; S4-6, For the first in the text entity set text entities If the x-coordinate of its geometric center Located in this extended collision zone Inside, the ordinate of its geometric center The coordinates of the nearest horizontal line segment are locked on both the top and bottom sides to generate the upper boundary. and lower boundary ; S4-7 Calculate the Y-axis projection domain of a single vertical line segment and Intersection length Its expression is: ; S4-8. A dynamic depth constraint threshold is constructed based on the adaptive fracture spacing, the standard deviation corresponding to the adaptive fracture spacing, the baseline visual ratio, and the font height fluctuation range corresponding to the baseline visual ratio. Its expression is:
[0028] in This is the dynamic depth constraint threshold; For adaptive break spacing; The standard deviation corresponding to the adaptive fracture spacing; Based on the baseline visual proportions; The range of font height fluctuation corresponding to the baseline visual proportion; S4-9, Judgment Is it greater than If so, the corresponding vertical line segment is determined to be a valid sidewall boundary, and the coordinates of the nearest valid sidewall boundary are locked on both the left and right sides of the vertical line segment to generate the corresponding left boundary. and right boundary Otherwise, the corresponding vertical line segment is determined to be an invalid boundary that is not a table structure, and it is not processed and skipped, that is, the above judgment logic is continued for the next candidate vertical line segment. S4-10. Determine whether the boundaries of the text entity in the four directions (up, down, left, and right) satisfy the intersection length constraint. If so, then... The four vertex coordinates of the virtual bounding box of the text entity are used to generate the virtual bounding box of the text entity; otherwise, it is determined that the ray has crossed the boundary in the direction that does not satisfy the intersection length constraint, and the boundary rollback mechanism is triggered. The intersection length constraint is not satisfied when the detection distance in that direction reaches the spatial infinity extreme value initialized by the algorithm. The out-of-bounds rollback mechanism includes: extracting the absolute extreme coordinates of the global geometric range of the current CAD drawing (i.e., the global upper boundary). Global lower boundary Global left boundary With the global right boundary The missing boundary in the direction where the ray crosses the boundary is forcibly assigned and constrained to the corresponding absolute extreme value coordinates, serving as the catch-all boundary for the incomplete table. This forces convergence and generates a logically closed virtual bounding box in extreme cases such as open tables or missing lines. .
[0029] In the specific implementation process, the projected coordinates are constructed from the vertical and horizontal boundary coordinates of all virtual bounding boxes. Then, through one-dimensional streaming mean clustering, the discrete virtual bounding box boundary coordinates are reconstructed into an orthogonal full-screen mathematical grid sequence with the origin at the top left corner. The specific methods include: S5-1. Project the vertical boundary coordinates of all virtual bounding boxes to obtain a one-dimensional horizontal projection coordinate set; project the horizontal boundary coordinates of all virtual bounding boxes to obtain a one-dimensional vertical projection coordinate set. With the first Virtual bounding box of each text entity For example, its one-dimensional horizontal projection coordinate set is denoted as The set of one-dimensional perpendicular projection coordinates is denoted as ; S5-2, Set of one-dimensional horizontal projection coordinates Perform a monotonically increasing sort to obtain an ordered sequence. ; S5-3, Based on Dynamic Stitching Operator The adaptive alignment tolerance is obtained by combining the network convergence weight coefficients, and its expression is:
[0030] in For adaptive alignment tolerance; These are the network convergence weights. S5-4, using ordered sequences Middle elements Establish an initial cluster family as the starting point; S5-5, Traversing the ordered sequence sequentially Middle elements The subsequent elements are then used to calculate the coordinate deviation, which is expressed as:
[0031] in This refers to coordinate deviation; For ordered sequences The Middle One element, ; For traversing elements The clustered summation at time, with an initial value of ; For traversing elements The number of cluster elements at time, with an initial value of 1; S5-6, if Then Absorb into the current cluster and sum the current clusters with The cluster sum is updated by adding the current number of cluster elements and incrementing the current number of cluster elements by 1. The next element is then traversed. S5-7, if If the current cluster is closed, calculate the quotient of the cluster sum and the number of cluster elements, and store it as the global standard column coordinate in the column grid sequence. Then, starting from the element where the coordinate deviation is currently calculated, the cluster sum and the number of cluster elements are reset, and the next round of clustering is started, until the ordered sequence has been traversed. Complete the column restructuring using all elements in the array. S5-8. Perform a monotonically descending sort on the one-dimensional perpendicular projection coordinate set, and then apply the sorting method to the ordered sequence. The same reconstruction method is used to reconstruct rows, extracting the arithmetic mean of each cluster to obtain the row grid sequence. This allows the discrete virtual bounding box boundary coordinates to be reconstructed into an orthogonal full-screen mathematical grid sequence with the origin at the top left corner.
[0032] Specifically, column grid sequence The vertical dividing lines (column physical boundaries) of the mesh are defined, and the row mesh sequence is defined. A horizontal dividing line (row physical boundary) is defined. These two sets of one-dimensional sequences are orthogonally combined (similar to the Cartesian product operation) to finally generate a two-dimensional table grid system that covers the entire screen and has the upper left corner as the logical coordinate origin (0,0).
[0033] In the specific implementation process, the physical boundary coordinates of the virtual bounding box are mapped to an orthogonal full-screen mathematical grid sequence for dimensionality reduction. Text entities with the same mapping relationship are grouped into the same logical cell. The basic two-dimensional logical coordinates and logical span of the text entities are calculated. The specific methods for obtaining the topological parameters of the corresponding logical cells include: S6-1, Extract the given physical boundary coordinates (e.g.) ), in the corresponding ordered grid sequence (such as Perform a heuristic extreme value approximation retrieval based on monotonicity in the following: S6-1-1. Traverse the orthogonal full-screen mathematical grid sequence in the forward direction, calculate the absolute difference between the coordinates of each physical boundary of the virtual bounding box and the coordinates of the current orthogonal full-screen mathematical grid node, and dynamically update and record the minimum difference and its corresponding sequence index. S6-1-2. In response to the current calculated absolute difference being greater than the sum of the historical minimum difference and a preset minimum tolerance, based on the monotonicity of the orthogonal full-screen mathematical grid sequence, it is determined that the optimal spatial matching point has been exceeded, triggering an early termination loop mechanism. The currently recorded sequence index is directly output as the grid node index obtained through mapping, thereby setting the two horizontal coordinates of the virtual bounding box... and Mapped to the starting column index respectively and End Column Index The two ordinates of the virtual bounding box and Mapped to the starting row index respectively and end row index ; S6-2, Based on the starting column index and End Column Index Calculate the logical column span of the corresponding text entity based on the starting row index. and end row index The logical line span of the corresponding text entity is calculated using the following expression:
[0034]
[0035] in For logical column span; For logical row span; S6-3. Constructing a basic two-dimensional logical coordinate system based on the starting row index and the starting column index. and based on this two-dimensional logical coordinate system Construct a unique hash value; S6-4. For multiple text entities mapped to the same hash value, determine that they have the same mapping relationship and classify them into the same logical cell; at the same time, dynamically compare and extract the maximum value of the span parameter of all related text entities under the basic two-dimensional logical coordinate (i.e., within the same logical cell), and use the largest span parameter within the same logical cell as the topology parameter of the logical cell; where the span parameter includes logical column span and logical row span. S6-5. Classify text content belonging to logical cells. The topology parameters of the logical cell are structured and bound together to output a structured logical data matrix.
[0036] Specifically, the introduction of "hash values" to replace direct coordinate comparison is primarily for deep optimization of system performance and algorithm time complexity: 1. Achieve ultra-fast retrieval and grouping: In real CAD drawings, a large table may contain tens of thousands of independent text entities. If the two-dimensional coordinates are traversed and compared every time, the algorithm's time complexity would be extremely high. By converting logical coordinates into unique hash values, the underlying hash table data structure can be directly utilized. This reduces the time complexity of text entity classification and retrieval.
[0037] 2. Solving the problem of multi-line text aggregation: Multiple lines of text often appear within the same table cell. Hash values can be used as "bucket" identifiers in memory to quickly aggregate scattered text entities belonging to the same cell and group them under the same hash key, avoiding complex nested loop comparisons.
[0038] In the specific implementation process, based on the ordered grid sequence With ordered grid sequence The coordinate difference between adjacent elements is mapped to generate the physical column width and physical row height parameters of the target spreadsheet; to ensure the closure of the global grid topology boundary of the target spreadsheet, the instantiation configuration of the full logical cell base map is performed in memory space in advance based on the dimensional characteristics of the ordered grid sequence; for empty cells without mapped valid data, memory placeholders are executed synchronously based on the physical column width and physical row height parameters.
[0039] In practice, specific methods for topologically merging cells in a target spreadsheet based on the fundamental two-dimensional logical coordinates and logical spans of text entities include: S7-1. Traverse the structured logical data matrix and extract the basic two-dimensional logical coordinates of each text entity. and its logical span parameters ;for or The text entity, with its basic two-dimensional logical coordinates As the reference anchor point, S7-2. Calculate and generate the absolute extreme boundary parameters defining the topology merging region: Calculate the termination row index of the topology merging region as follows: The index of the terminating column of the topology merge region is calculated as follows. ; S7-3. Based on the baseline anchor point and the corresponding termination row index and termination column index of the topology merging area, the NPOI underlying application interface is called to construct the corresponding topology merging area for cells spanning rows and columns, and the synchronous rendering instruction of the outer border of the topology merging area is executed to complete the topology merging of cells in the target spreadsheet.
[0040] In the specific implementation process, the methods for dimensionality reduction sorting and concatenation of text entities within the same logical cell, combined with the topological parameters of the corresponding logical cell, filling the logical coordinates of the target spreadsheet, and serializing the output include: S8-1. Constructing a local text subset based on multiple discrete text entities mapped to the same basic two-dimensional logical coordinates. and retrieve the local text subset. The geometric center coordinates of each text entity within the text; S8-2, using the geometric center coordinates of the text entity For example, taking the primary sorting key as the geometric center coordinate of the Y-axis. Furthermore, the keys are arranged in descending order, with the X-axis geometric center coordinates used for sorting. Furthermore, by arranging them in an increasing direction, local text subsets are mapped to a one-dimensional linear sequence; S8-3. To eliminate the data redundancy defects caused by overlapping elements in the source drawing, perform a deduplication filtering operation based on hash matching on the text string in the one-dimensional linear sequence to obtain the filtered one-dimensional linear sequence. S8-4. Text content in the filtered one-dimensional linear sequence Perform character concatenation and recombination; S8-5. Perform adaptive data type conversion on the concatenated and recombined text content: If the concatenated and recombined text content matches the preset pure numeric rule, then the concatenated and recombined text content will be adaptively converted to numeric floating-point data; otherwise, the string data format will be maintained. S8-6. Fill the logical coordinates of the target spreadsheet with the text content that has undergone adaptive data type conversion, and perform cell merging configuration on the above logical coordinates according to its corresponding row and column topology attributes, and then serialize and output it as a standardized structured spreadsheet file.
[0041] Specifically, the text content is "filled" into the logical coordinates of the target spreadsheet as actual data; while the row and column topology attributes are used as structural instructions to "configure" or "reshape" the merged state of the logical coordinates (i.e., to perform cell merging operations).
[0042] In the specific implementation process, a system based on the unstructured CAD table extraction method is also provided, which includes: The first module is used to receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and to perform heterogeneous entity deconstruction and spatial feature vectorization, and output the line segment entity set and the text entity set. The second module is used to dynamically extract the baseline visual proportion of the drawing based on the statistical frequency distribution principle for the text entity set, and then obtain the font height fluctuation range in the current drawing; for the line segment entity set, it adaptively calculates the global tolerance of physical break and trimming error based on the spatial nearest neighbor statistical algorithm, and then performs angle feature filtering to obtain the horizontal line set and the vertical line set. The third module is used to initiate orthogonal rays in the set of horizontal and vertical lines with the geometric center of each text entity as the probe origin, and then capture the boundary of each text entity by combining the font height fluctuation range, and generate a logically closed virtual bounding box for the text entity. The fourth module is used to construct projected coordinates for the vertical and horizontal boundary coordinates of all virtual bounding boxes, and then reconstruct the discrete virtual bounding box boundary coordinates into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner through one-dimensional streaming mean clustering. The fifth module is used to perform dimensionality reduction mapping of the physical boundary coordinates of the virtual bounding box to the orthogonal full-screen mathematical grid sequence, group text entities with the same mapping relationship into the same logical cell, calculate the basic two-dimensional logical coordinates and logical span of the text entity, and obtain the topological parameters of the corresponding logical cell. The sixth module is used to generate the physical column width and row height parameters of the target spreadsheet based on the coordinate difference mapping of adjacent elements in the orthogonal full-screen mathematical grid sequence; and to perform topological merging of the cells of the target spreadsheet based on the basic two-dimensional logical coordinates and logical span of the text entity. The seventh module is used to sort and concatenate text entities within the same logical cell in a dimensionality reduction manner, fill them into the logical coordinates of the target spreadsheet with the topological parameters of the corresponding logical cell, and serialize and output them to complete the extraction of unstructured CAD tables.
[0043] In one embodiment of the present invention, in order to more intuitively present the execution process of the core algorithm of the present invention and its strong tolerance to the background noise of manual drawing, this embodiment takes a typical unstructured scattered table in a CAD drawing with "two rows and two columns in a local area, including merged headers and fragmented text" as an example, inputs real coordinate data and deduces the whole process of the algorithm.
[0044] like Figure 3 As shown, assume the target table in the CAD drawing is composed of the following discrete entities and has obvious drawing defects. Detailed data is as follows: Text entities ( ): T1 (Header Content): "Material Details", Height ,coordinate ; T2 (bottom left, content 1): "Steel", height ,coordinate ; T3 (bottom left, content two): "Material", height ,coordinate (Note: This refers to the common phenomenon of single-line text being interrupted in engineering projects.) T4 (bottom right): "Q235", height ,coordinate .
[0045] Line segment entity ( ): Horizon (Top line): absolute level; Horizon (Midline: There is a slight tilt of 0.1° due to hand tremor at the left endpoint (10.0, 80.0) and the right endpoint (90.0, 79.8). vertical line (Left side line): ; vertical line (Middle vertical line): Upper endpoint (50.0, 80.0), lower endpoint (50.0, 48.0) (There is a cross-line burr that crosses the bottom line by 2.0 length). vertical line (Right line): Upper endpoint (90.0, 100.0), lower endpoint (90.0, 50.5) (not connected to the bottom line, there is a break of 0.5 length).
[0046] This invention employs a recursive destructuring (physical decomposition) algorithm to penetrate potential block parameters within the outer layer of the table and extract the underlying entities. It also calculates the geometric center of the outer box for each text entity. This serves as the absolute origin for subsequent ray projections. For example: Geometric center of T1 ; Geometric center of T2 .
[0047] Next, the present invention automatically "reads" the inherent features of the drawings, rejecting fixed empirical values: 1. Height and Breakpoint Extraction: The system calculates the mode of the text height. Extract the distance between the endpoints of the line segments and discover the spatial nearest neighbor statistics. Breakage features at the location, median extracted .
[0048] 2. Generate dynamic tolerance The system incorporates the optimal weight coefficients. Calculate the dynamic stitching operator .
[0049] 3. Angle Filtering and Orthogonalization: The system calculates the inclination angle of each line segment based on... Criterion for generating parallelism tolerance Due to the center line The tilt angle (0.1°) is less than 0.1°, successfully absorbing it into the horizontal line set. It is perfectly immune to the local coordinate system tilt caused by hand tremors.
[0050] Then, using the header text T1 (center) Taking (e.g.) as an example, initiate an orthogonal optimization: 1. Suture corner breakage: When ray is emitted to the right, due to the right side line... The Y-axis range only extends to 50.5, which originally could not wrap around T1 ( However, the extended collision region was invoked. (That is, the range was expanded by 1.5), and the ray was successfully captured. As the right boundary .
[0051] 2. Penetrating interference burrs: Vertical lines in the judgment. When calculating the intersection length with the upper and lower boundaries, calculate the intersection length with the upper and lower boundaries. Calculate the anti-penetration threshold. Because of 30.0 1.5, System Judgment It is an effective sidewall; conversely, the 2.0-length burr protruding downwards to 48.0 is completely ignored because it does not meet the conditions.
[0052] Ultimately, a logically closed virtual bounding box was successfully synthesized for T1. (Four coordinates: top 100.0, bottom 79.8, left 10.0, right 90.0).
[0053] Extract the physical X-boundaries of all virtual bounding boxes and monotonically sort them to obtain a jittery coordinate sequence: .
[0054] Introducing grid convergence coefficients And generate alignment tolerance by combining the safety bottom line threshold. .
[0055] Clustering begins: 1. The initial cluster is set to 10.0. When iterating to 10.1, the difference is 0.1 ≤ 1.0, so it is absorbed into the same cluster, generating the current arithmetic mean of 10.05.
[0056] 2. The cursor moves to 50.0, the difference 39.95 > 1.0, a fault occurs! The system determines that the previous cluster is closed, outputs the absolute logical grid coordinates 10.05, and opens a new cluster with 50.0.
[0057] After a one-dimensional streaming traversal, the originally jagged six physical boundaries were perfectly reduced in dimension and reconstructed into three absolutely perpendicular mathematical logic column grids: .
[0058] Then, using the standard grid sequence Reverse the process of merging cell properties: 1. Heuristic extremum approximation mapping: Input the left physical boundary 10.0 of the header text T1. The search revealed that the absolute difference between the current value and the first node (10.05) was 0.05. However, when traversing to the second node (50.05), the difference surged to 40.05. Detecting this "bottoming out" divergence, the system immediately triggered an early termination command, instantly locking the nearest grid node at 10.05 (column index). Similarly, the right boundary 90.0 is quickly matched to grid node 90.05 (column index). ).
[0059] 2. Topological span inversion: Perform basic interpolation directly on the obtained grid index to obtain the logical column span. This determines that the cell containing the text T spans two columns. This mechanism completely eliminates the need for complex graphic collision detection and strict ascending / descending order restrictions, greatly reducing invalid calculation branches.
[0060] Continuing with this method, it was detected that text T2 "steel" (coordinates 20.0, 65.0) and text T3 "material" (coordinates 40.0, 65.0) were mapped to the same logical cell. This triggered a strict two-dimensional dimensionality reduction sort: since the Y-coordinates were identical, the X-coordinates were compared, and it was determined that 20.0 < 40.0, therefore "steel" was placed before "material". Simultaneously, the underlying hash deduplication mechanism automatically filtered out any possible in-situ overlapping copies, ultimately concatenating and outputting the clean semantic string "steel material".
[0061] In the export stage, the grid sequence is used in advance. and The spacing is initialized with all empty cells and border closure placeholders to prevent style loss caused by merging. Then, the target text is type-detected (if a pure numeric value is detected, an implicit conversion will be triggered; however, in this embodiment, "Q235" in T4 contains letters, so it will be adaptively preserved as a string). Finally, by calling the NOPI underlying API to generate the merged area, the header "Material Details" is accurately filled into the merged cells spanning two columns, outputting a 100% topologically correct, engineering-grade structured spreadsheet.
[0062] In summary, this invention abandons the hard-coded tolerance of traditional fixed values or arithmetic means, and instead extracts the mode of the total text height and the median of the break distance under spatial nearest neighbor statistics, combined with the normal distribution. The principle is to automatically derive dynamic stitching operators that highly fit the inherent geometric features of the current drawing. Parallelism tolerance Based on this, a one-dimensional streaming mean clustering algorithm is used to perform noise reduction and normalization on the coordinates of discrete line segments with physical jitter, forcing the free endpoints within the adaptive tolerance range to converge into an absolutely orthogonal logical grid sequence, thus laying a rigorous mathematical foundation for subsequent topology reconstruction.
[0063] Unlike traditional closure detection methods that require line segments to be strictly connected end-to-end to form polygons, this invention pioneers an orthogonal ray projection optimization algorithm with the geometric center of the text as the spatial anchor point. This algorithm combines a dynamic depth constraint threshold (to adaptively filter out redundant line segments) and a dynamically expanded collision interval (to forcibly mend broken gaps), dynamically locking the nearest valid physical boundary in four orthogonal directions, thereby generating an absolutely closed "virtual bounding box" for discrete text entities. Subsequently, the system establishes a spatial mapping relationship between the two-dimensional physical bounding box and a one-dimensional orthogonal logical grid sequence. Through low-complexity grid index difference operations (i.e., logical span calculation of row and column indices), it accurately inversely calculates the attributes of merged cells with arbitrary complex spans, achieving lossless restoration of the topology.
[0064] In the initial stage of spatial analysis, this invention pre-decouples and dimensionality-reduced the massive global discrete line segments into absolute sets of horizontal and vertical lines using a filtering mechanism based on the statistical peak angle, instantly reducing the subsequent search space by more than half. When mapping virtual bounding box coordinates to a standard grid sequence, this invention abandons the time-consuming global blind linear traversal and does not employ the traditional binary search constrained by strict monotonic sorting rules; instead, it innovatively introduces a heuristic extreme value approximation search mechanism based on sequence monotonicity. This mechanism cleverly utilizes the "bottoming out" characteristic of absolute differences during spatial coordinate mapping, triggering an early termination instruction once a diverging trend in the difference is detected. This strategy not only achieves immune compatibility with different monotonic directions of the X and Y axes but also significantly reduces invalid computational branches in practical engineering applications, achieving limit alignment of coordinate nodes. This deep coupling of physical dimension decoupling and extreme value approximation acceleration strategy enables this invention to maintain top-tier analytical performance and extremely low memory overhead even when dealing with extremely dense, extra-large CAD engineering drawings.
Claims
1. A method for extracting unstructured CAD tables, characterized in that, Includes the following steps: Receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and perform heterogeneous entity deconstruction and spatial feature vectorization, outputting a set of line segment entities and a set of text entities; For a set of text entities, the baseline visual proportion of the drawing is dynamically extracted based on the statistical frequency distribution principle, thereby obtaining the fluctuation range of font height in the current drawing; For the set of line segments, the global tolerance of physical line breakage and trimming error is adaptively calculated based on the spatial nearest neighbor statistical algorithm, and then the angle feature filtering is performed to obtain the set of horizontal lines and the set of vertical lines. Using the geometric center of each text entity as the origin of the probe, orthogonal rays are initiated in the set of horizontal and vertical lines. Then, the boundaries of each text entity are captured by combining the font height fluctuation range, and a logically closed virtual bounding box is generated for the text entity. Projected coordinates are constructed from the vertical and horizontal boundary coordinates of all virtual bounding boxes. Then, the discrete virtual bounding box boundary coordinates are reconstructed into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner by one-dimensional streaming mean clustering. The physical boundary coordinates of the virtual bounding box are mapped to the orthogonal full-screen mathematical grid sequence for dimensionality reduction. Text entities with the same mapping relationship are grouped into the same logical cell, and the basic two-dimensional logical coordinates and logical span of the text entity are calculated. At the same time, the topological parameters of the corresponding logical cell are obtained. The physical column width and row height parameters of the target spreadsheet are generated based on the coordinate difference mapping of adjacent elements in an orthogonal full-screen mathematical grid sequence. Topological merging of cells in the target spreadsheet based on the underlying two-dimensional logical coordinates and logical span of text entities; Text entities within the same logical cell are sorted and concatenated in a dimensionality reduction manner. The topological parameters of the corresponding logical cells are then used to fill the logical coordinates of the target spreadsheet and serialized for output, thus completing the extraction of unstructured CAD tables. Specific methods for dimensionality reduction, sorting, and concatenating text entities within the same logical cell include: A local text subset is constructed based on multiple discrete text entities mapped to the same basic two-dimensional logical coordinates, and the geometric center coordinates of each text entity within the local text subset are retrieved. Using the primary sorting key as the geometric center coordinate of the Y-axis and arranged in a decreasing direction, and the secondary sorting key as the geometric center coordinate of the X-axis and arranged in an increasing direction, a local text subset is mapped into a one-dimensional linear sequence. Perform a hash-match-based deduplication filtering operation on the text string within a one-dimensional linear sequence to obtain the filtered one-dimensional linear sequence. Perform character concatenation and recombination on the text content in the filtered one-dimensional linear sequence.
2. The method for extracting unstructured CAD tables according to claim 1, characterized in that, The specific methods for receiving unstructured CAD table entity objects selected by the user in the target CAD drawing and performing heterogeneous entity deconstruction and spatial feature vectorization include: Receive the unstructured CAD table entity object to be exported selected by the user in the target CAD drawing, and extract the underlying type identifier of the unstructured CAD table entity object; For unstructured CAD table entity objects whose underlying type is identified as composite encapsulated entities, the CAD underlying decomposition interface is called to perform physical decomposition operations on the composite encapsulated entities, generating a discrete set of sub-entities. The physical decomposition algorithm is recursively called with each sub-entity in the sub-entity set as input parameter to obtain the basic primitive features contained in the composite encapsulated entity; If the basic primitive feature is a straight line, it is stored in a preset set of line segment entities; if the basic primitive feature is a text entity, the height and content of each text entity are extracted, the geometric center coordinates of the bounding box of the text entity are calculated, and the text entity is stored in a preset set of text entities. For unstructured CAD table entity objects whose underlying type is a standard polyline, traverse their vertex sequence and extract the sub-segment type formed by adjacent vertices; when the sub-segment type is a line segment, extract the spatial coordinate data of the current vertex and the next adjacent point, convert them into independent basic line segment features and store them in the preset line segment entity set.
3. The method for extracting unstructured CAD tables according to claim 1, characterized in that, The specific methods for dynamically extracting the baseline visual scale of a drawing based on the statistical frequency distribution principle, and then obtaining the fluctuation range of font height in the current drawing, include: Extract the height of all text entities in the text entity set, and then construct a global text height sample sequence; Calculate the initial average height of the global text height sample sequence, and set 5% to 10% of the initial average height as the height discretization step size; Based on the height discretization step size, the global text height sample sequence is mapped to the corresponding height intervals, the entity frequency in each height interval is counted, and a height frequency histogram is generated. Peak retrieval is performed in the height frequency histogram to identify the highest frequency maximum peak interval, and the central tendency feature value of the samples in the maximum peak interval is extracted and used as the global baseline visual scale of the drawing. Calculate the standard deviation of the sample deviation from the global baseline visual proportion within the maximum peak interval, and use it as the font height fluctuation range in the current drawing.
4. The method for extracting unstructured CAD tables according to claim 1, characterized in that, The specific methods for adaptively calculating the global tolerance of physical line breakage and trimming errors based on the spatial nearest neighbor statistical algorithm, and then performing angular feature filtering to obtain the set of horizontal lines and the set of vertical lines include: Calculate the inclination angle of each line segment in the line segment entity set, and then construct the horizontal angle density histogram and the vertical angle density histogram respectively; The angle of the highest main peak in the horizontal angle density histogram is used as the benchmark horizontal angle. And calculate the standard deviation of the horizontal angles in the horizontal angle density histogram. ; The angle of the highest main peak in the vertical angle density histogram is used as the reference vertical angle. And calculate the standard deviation of the vertical angle in the vertical angle density histogram. ; Based on the Laida criterion, a horizontal parallelism tolerance is generated. and vertical parallelism tolerance ;in To prevent the tolerance from reaching a minimum of zero, the horizontal parallelism tolerance and the vertical parallelism tolerance constitute the global tolerance. By applying angular feature filtering to the line segment entity set using global tolerance, we obtain the horizontal line set and the vertical line set. The corresponding expression is: in A set of horizontal lines; A set of vertical lines; representing; Represents a set of line segment entities The first in A single line segment entity; Line segment entity The angle of inclination.
5. The method for extracting unstructured CAD tables according to claim 1, characterized in that, Using the geometric center of each text entity as the probe origin, orthogonal rays are initiated from the sets of horizontal and vertical lines. Then, the boundaries of each text entity are captured by combining the font height fluctuation range. Specific methods for generating logically closed virtual bounding boxes for text entities include: Traverse each independent line segment in the line segment entity set, calculate the Euclidean distance from the endpoint of each line segment to the endpoints of other non-collinear line segments in its spatial neighborhood; after removing completely closed points with zero Euclidean distance, construct a sample set of endpoint spacing that characterizes the physical break scale of the drawing. The endpoint spacing sample set is sorted in ascending order, the median in the middle of the sequence is extracted, and this median is used as the adaptive break spacing for the overall broken line defect scale of the current drawing; the standard deviation of the sequence is obtained based on the median in the middle of the sequence, which is the standard deviation corresponding to the adaptive break spacing. The dynamic stitching operator is obtained by weighting the baseline visual scale of the drawing with the adaptive fracture spacing. ; Traverse the set of text entities, using the geometric center of each text entity as the probe origin, and initiate orthogonal rays in the sets of horizontal and vertical lines to obtain the X-axis projection domain of a single horizontal line segment in the set of horizontal lines. Y-axis projection domain of a single vertical line segment in the set of vertical lines ;in and These are the minimum and maximum coordinate values of the X-axis projection domain of a single horizontal line segment, respectively. and These are the minimum and maximum coordinate values of the Y-axis projection domain for a single vertical line segment, respectively. X-axis projection domain of a single horizontal line segment Based on this, a dynamic stitching operator is introduced. Construct extended collision range ; For the first in the text entity set text entities If the x-coordinate of its geometric center Located in this extended collision zone Inside, the ordinate of its geometric center The coordinates of the nearest horizontal line segment are locked on both the top and bottom sides to generate the upper boundary. and lower boundary ; Calculate the Y-axis projection domain of a single vertical line segment and Intersection length ; A dynamic depth constraint threshold is constructed based on the adaptive fracture spacing, the standard deviation corresponding to the adaptive fracture spacing, the baseline visual ratio, and the font height fluctuation range corresponding to the baseline visual ratio. Its expression is as follows: in This is the dynamic depth constraint threshold; For adaptive break spacing; The standard deviation corresponding to the adaptive fracture spacing; Based on the baseline visual proportions; The range of font height fluctuation corresponding to the baseline visual proportion; judge Is it greater than If so, the corresponding vertical line segment is determined to be a valid sidewall boundary, and the coordinates of the nearest valid sidewall boundary are locked on both the left and right sides of the vertical line segment to generate the corresponding left boundary. and right boundary Otherwise, the corresponding vertical line segment is determined to be an invalid boundary that is not a table structure, and it is not processed and is skipped. Determine whether the boundaries of the text entity in the four directions (up, down, left, right) satisfy the intersection length constraint. If so, then... The four vertex coordinates of the virtual bounding box of the text entity are used to generate the virtual bounding box of the text entity; otherwise, it is determined that the ray has crossed the boundary in the direction that does not satisfy the intersection length constraint, and the boundary rollback mechanism is triggered. The boundary rollback mechanism includes: extracting the absolute extreme coordinates of the global geometric range of the current CAD drawing, forcibly assigning and constraining the missing boundary of the ray crossing direction to the corresponding absolute extreme coordinates, serving as the catch-all boundary of the incomplete table, thereby forcibly converging and generating a logically closed virtual bounding box in extreme cases of open tables or missing lines.
6. The method for extracting unstructured CAD tables according to claim 5, characterized in that, The specific method for constructing projected coordinates from the vertical and horizontal boundary coordinates of all virtual bounding boxes, and then reconstructing the discrete virtual bounding box boundary coordinates into an orthogonal full-screen mathematical grid sequence with a starting origin at the top left corner using one-dimensional streaming mean clustering, includes: Projecting the vertical boundary coordinates of all virtual bounding boxes yields a one-dimensional horizontal projection coordinate set; projecting the horizontal boundary coordinates of all virtual bounding boxes yields a one-dimensional vertical projection coordinate set. Perform a monotonically increasing sort on the set of one-dimensional horizontal projected coordinates to obtain an ordered sequence. ; Based on the dynamic stitching operator and the network convergence weight coefficients, the adaptive alignment tolerance is obtained, and its expression is as follows: in For adaptive alignment tolerance; These are the network convergence weights. With ordered sequence Middle elements Establish an initial cluster family as the starting point; Traverse the ordered sequence sequentially Middle elements The subsequent elements are then used to calculate the coordinate deviation, which is expressed as: in This refers to coordinate deviation; For ordered sequences The Middle One element, ; For traversing elements The clustered summation at time, with an initial value of ; For traversing elements The number of cluster elements at time, with an initial value of 1; like Then Absorb into the current cluster and sum the current clusters with The cluster sum is updated by adding the current number of cluster elements and incrementing the current number of cluster elements by 1. The next element is then traversed. like If the current cluster is closed, calculate the quotient of the cluster sum and the number of cluster elements, and store it as the global standard column coordinate in the column grid sequence. Then, starting from the element where the coordinate deviation is currently calculated, the cluster sum and the number of cluster elements are reset, and the next round of clustering is started, until the ordered sequence has been traversed. Complete the column restructuring using all elements in the array. Perform a monotonically descending sort on the set of one-dimensional orthogonal projection coordinates, and then apply the sorted sequence... The same reconstruction method is used to reconstruct rows, resulting in a row grid sequence. This allows the discrete virtual bounding box boundary coordinates to be reconstructed into an orthogonal full-screen mathematical grid sequence with the origin at the top left corner.
7. The method for extracting unstructured CAD tables according to claim 6, characterized in that, The specific methods for performing dimensionality reduction mapping of the physical boundary coordinates of the virtual bounding box to an orthogonal full-screen mathematical grid sequence, grouping text entities with the same mapping relationship into the same logical cell, and calculating the basic two-dimensional logical coordinates and logical span of the text entities, while obtaining the topological parameters of the corresponding logical cells, include: Traverse the orthogonal full-screen mathematical grid sequence in a forward direction, calculate the absolute difference between the physical boundary coordinates of the virtual bounding box and the coordinates of the current orthogonal full-screen mathematical grid node, and dynamically update and record the minimum difference and its corresponding sequence index. If the calculated absolute difference is greater than the sum of the historical minimum difference and a preset minimum tolerance, then based on the monotonicity of the orthogonal full-screen mathematical grid sequence, it is determined that the optimal spatial matching point has been exceeded, triggering an early termination loop mechanism. The currently recorded sequence index is then directly output as the grid node index obtained through mapping, thereby mapping the two horizontal coordinates of the virtual bounding box. and Mapped to the starting column index respectively and End Column Index The two ordinates of the virtual bounding box and Mapped to the starting row index respectively and end row index ; Based on the index of the start column and End Column Index Calculate the logical column span of the corresponding text entity based on the starting row index. and end row index The logical line span of the corresponding text entity is calculated using the following expression: in For logical column span; For logical row span; A basic two-dimensional logical coordinate system is constructed based on the starting row index and the starting column index. and based on this two-dimensional logical coordinate system Construct a unique hash value; For multiple text entities mapped to the same hash value, they are determined to have the same mapping relationship and are grouped into the same logical cell. The largest span parameter within the same logical cell is used as the topology parameter of that logical cell; where the span parameter includes the logical column span and the logical row span. The text content belonging to a logical cell is structurally bound to the topology parameters of that logical cell, thereby outputting a structured logical data matrix.
8. The method for extracting unstructured CAD tables according to claim 7, characterized in that, Specific methods for topologically merging cells in a target spreadsheet based on the fundamental two-dimensional logical coordinates and logical spans of text entities include: Traverse the structured logical data matrix to extract the basic two-dimensional logical coordinates of each text entity. and its logical span parameters ;for or The text entity, with its basic two-dimensional logical coordinates Using the baseline anchor point, calculate the terminating row index of the topology merge region as follows: The index of the terminating column of the topology merge region is calculated as follows. ; Based on the baseline anchor point and the corresponding termination row index and termination column index of the topology merging region, the NPOI underlying application interface is called to construct the corresponding topology merging region for cells spanning rows and columns, and the synchronous rendering instruction for the outer border of the topology merging region is executed, thereby completing the topology merging of the cells in the target spreadsheet.
9. The method for extracting unstructured CAD tables according to claim 8, characterized in that, The specific methods for filling the logical coordinates of the target spreadsheet with the topology parameters of the corresponding logical cells and then serializing and outputting them include: Perform adaptive data type conversion on the concatenated and recombined text content: if the concatenated and recombined text content matches the preset pure numeric rule, then the concatenated and recombined text content will be adaptively converted to numeric floating-point data; otherwise, the string data format will be maintained. The text content, after adaptive data type conversion, is filled into the logical coordinates of the target spreadsheet. Based on its corresponding row and column spanning topology attributes, the logical coordinates are configured to merge cells, and then serialized and output as a standardized structured spreadsheet file.
10. A system based on the unstructured CAD table extraction method according to any one of claims 1 to 9, characterized in that, include: The first module is used to receive the unstructured CAD table entity objects to be exported selected by the user in the target CAD drawing, and to perform heterogeneous entity deconstruction and spatial feature vectorization, and output the line segment entity set and the text entity set. The second module is used to dynamically extract the baseline visual proportion of the drawing based on the statistical frequency distribution principle for the text entity set, and then obtain the font height fluctuation range in the current drawing; for the line segment entity set, it adaptively calculates the global tolerance of physical break and trimming error based on the spatial nearest neighbor statistical algorithm, and then performs angle feature filtering to obtain the horizontal line set and the vertical line set. The third module is used to initiate orthogonal rays in the set of horizontal and vertical lines with the geometric center of each text entity as the probe origin, and then capture the boundary of each text entity by combining the font height fluctuation range, and generate a logically closed virtual bounding box for the text entity. The fourth module is used to construct projected coordinates for the vertical and horizontal boundary coordinates of all virtual bounding boxes, and then reconstruct the discrete virtual bounding box boundary coordinates into an orthogonal full-screen mathematical grid sequence with the starting origin at the top left corner through one-dimensional streaming mean clustering. The fifth module is used to perform dimensionality reduction mapping of the physical boundary coordinates of the virtual bounding box to the orthogonal full-screen mathematical grid sequence, group text entities with the same mapping relationship into the same logical cell, calculate the basic two-dimensional logical coordinates and logical span of the text entity, and obtain the topological parameters of the corresponding logical cell. The sixth module is used to generate the physical column width and row height parameters of the target spreadsheet based on the coordinate difference mapping of adjacent elements in the orthogonal full-screen mathematical grid sequence; Topological merging of cells in the target spreadsheet based on the underlying two-dimensional logical coordinates and logical span of text entities; The seventh module is used to sort and concatenate text entities within the same logical cell in a dimensionality reduction manner, fill them into the logical coordinates of the target spreadsheet with the topological parameters of the corresponding logical cell, and serialize and output them to complete the extraction of unstructured CAD tables.
Citation Information
Patent Citations
Methods and systems for recognizing tables in CAD drawings
CN111553187B
List compilation method and device based on drawing table, equipment and storage medium
CN116933350A
Spatial position-based method for reconstructing text information in CAD electronic data
CN105808511A
Automatic extraction method for CAD (Computer Aided Design) detail list based on Autolisp
CN121259855A