Building CAD multi-modal extraction and asset data generation method and system

By parsing text annotations and visual guidance information in architectural CAD drawings, strong and weak constraint relationships are established, which solves the problems of dense overlapping of components and ambiguous annotations, improves the accuracy and consistency of asset field values, and reduces the workload of manual verification.

CN122065360APending Publication Date: 2026-05-19CHINA RAILWAY 14TH BUREAU GRP CO LTD ELECTRICAL SERVICE ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY 14TH BUREAU GRP CO LTD ELECTRICAL SERVICE ENG CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In architectural CAD drawings, dense overlapping of components, crowded annotations, and ambiguous annotation directions make it difficult for text annotations to uniquely correspond to components, resulting in incorrect and inconsistent asset field values, and increasing manual verification and rework costs.

Method used

By acquiring architectural CAD files, parsing text annotations, component geometry, and visual guidance element information, establishing strong constraint binding relationships and weak constraint candidate relationships, and using visual guidance information and competitive arbitration rules to determine the component to which the text annotations belong, asset field values ​​are generated.

Benefits of technology

It improves the robustness and consistency of automatic attribution determination under complex drawing conditions, reduces the probability of mis-binding and omission binding, reduces the workload of manual review and correction, and adapts to the large-scale processing needs of engineering asset data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065360A_ABST
    Figure CN122065360A_ABST
Patent Text Reader

Abstract

The invention provides a building CAD multi-modal extraction and asset data generation method and system, and relates to the technical field of computer aided design, the method comprises the following steps: obtaining and analyzing a building CAD file to obtain text labeling information, component geometric information and visual guidance primitive information; generating a component candidate set based on the component geometric information; determining a source text label and a target area pointed by the source text label by utilizing association between the visual guidance primitive and the text label, determining a target component candidate in the candidate set, and establishing a strong constraint binding relationship; and constructing a weak constraint candidate relationship for the source text annotations which do not form strong constraint binding, determining attribution component candidates through competitive arbitration, and mapping asset attribute items obtained by analyzing the source text annotations into predefined asset information structures to generate asset field values. The method is suitable for component dense overlapping and pointing ambiguity scenes, and the accuracy and stability of affiliation determination are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer-aided design technology, and more specifically, to a method and system for multimodal extraction and asset data generation in architectural CAD. Background Technology

[0002] In the design and operation of building engineering projects, two-dimensional architectural CAD drawings are still widely used to express component outlines, pipeline routes, and their material, specifications, elevation, and other attribute information. To support asset ledger establishment, operation and maintenance inspections, renovation assessments, and digital delivery, the industry typically needs to establish a correspondence between text annotations in CAD drawings and component geometric objects, and structure the annotation content into calculable and searchable asset field values. Current technologies often employ parsing and rule-based processing of CAD files, combined with heuristic criteria such as layer semantics, object type, spatial proximity, and bounding box overlap, to assign annotations to specific component or pipeline objects.

[0003] However, in real engineering drawings, issues such as densely overlapping components, crowded annotations due to zooming in on partial views, non-standard expressions of leader lines / dimension lines, and ambiguous annotation directions are very common, making it difficult to uniquely correspond text annotations to components. Relying solely on spatial proximity or static rules can easily lead to misbinding, omissions, or unstable attribution, resulting in incorrect asset field values, poor asset consistency, and increased costs for subsequent manual verification and rework. Therefore, in the process of parsing architectural CAD drawings and generating asset data, facing complex situations such as densely overlapping components, ambiguous annotation directions, and difficulty in uniquely corresponding annotations to components, improving the automatic attribution determination mechanism between annotations and components to enhance the accuracy and stability of attribution determination, thereby improving the correctness and consistency of asset field value generation, has become an urgent technical problem to be solved. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this application provides a method and system for multimodal extraction and asset data generation in architectural CAD.

[0005] Firstly, this application provides a method for multimodal extraction and asset data generation in architectural CAD, including:

[0006] The architectural CAD file is acquired and parsed to obtain text annotation information, component geometric information, and visual guidance element information.

[0007] Based on the geometric information of the components, a candidate set of components is generated;

[0008] Based on the visual guidance primitive information and the text annotation information, visual guidance primitives are determined from the visual guidance primitive information, source text annotations associated with the visual guidance primitives and target areas pointed to by the visual guidance primitives are determined, and target component candidates corresponding to the target areas are determined in the component candidate set, and a strong constraint binding relationship is established between the source text annotations and the target component candidates.

[0009] For source text annotations that do not have a strong constraint binding relationship or whose strong constraint binding relationship cannot be uniquely determined, a weak constraint candidate relationship is established between the source text annotation and the component candidate based on a preset association criterion.

[0010] Based on the strong constraint binding relationship and the weak constraint candidate relationship, the candidate attribution component corresponding to the source text annotation is determined according to the competitive arbitration rule, and the asset attribute items obtained by parsing the source text annotation are mapped to a predefined asset information structure to generate asset field values.

[0011] Optionally, determining the source text annotation associated with the visual guidance primitive and the target area pointed to by the visual guidance primitive includes:

[0012] Visual guidance elements are determined from the visual guidance element information, and source text annotations are determined based on the association between the visual guidance elements and the text annotation information; the target area is determined based on the pointing end of the visual guidance elements.

[0013] Optionally, the competitive arbitration rules include:

[0014] When the strong constraint binding relationship corresponding to the source text annotation exists and is unique, the belonging component candidate is determined from the target component candidates corresponding to the strong constraint binding relationship;

[0015] When the source text annotation does not establish a strong constraint binding relationship, or the strong constraint binding relationship corresponds to multiple different target component candidates and cannot be uniquely determined, the belonging component candidate is determined based on the weak constraint candidate relationship.

[0016] Optionally, determining the visual guidance primitive from the visual guidance primitive information includes:

[0017] The entity object set of the architectural CAD file is traversed, and a guiding candidate set is constructed based on object type identifier, style parameters, geometric parameters, and topological connection relationships between entity objects;

[0018] Based on the object type identifier and style parameters, entity objects that meet the preset guidance object identification conditions are selected from the guidance candidate set as the first type of guidance candidate;

[0019] Based on the endpoint features and topological connection features of basic geometric primitives, combinations of geometric primitives that meet preset geometric topological guidance conditions are selected from the guidance candidate set as the second type of guidance candidate;

[0020] Geometric normalization and attribute normalization are performed on the first type of guide candidate and the second type of guide candidate to obtain visual guide primitives.

[0021] Optionally, the preset geometric topology guiding conditions include:

[0022] The first end of the basic geometric primitive is connected to an indicator symbol entity in a preset set of indicator symbols, the indicator symbol entity being used to represent a pointing arrow or an indicator end mark;

[0023] The second end of the basic geometric primitive is located within a preset neighborhood of the text annotation object, and the preset neighborhood is determined by the bounding box of the text annotation object and the expansion distance corresponding to the bounding box;

[0024] The directional characteristics of the basic geometric primitives are consistent with the pointing direction of the indicator symbol entity within a preset angle range;

[0025] Under the conditions of the preset indicator symbol connection condition, the preset neighborhood condition, and the direction consistency condition, the basic geometric primitive and the indicator symbol entities connected to it are combined and reconstructed into a single visual guidance primitive.

[0026] Optionally, the preset set of indicator symbols is determined in the following way:

[0027] From the entity objects in the architectural CAD file, select entity objects that meet the preset indicator symbol determination rules, and determine the selected entity objects as the indicator symbol entities; wherein, the preset indicator symbol determination rules include at least one of the following:

[0028] The entity object is a block reference object, and its block name belongs to the preset arrow block name set;

[0029] The entity object is a closed graphic object with preset fill style parameters, and its outer contour contains sharp corner vertices to represent the pointing end;

[0030] The entity object is a short line segment primitive with a preset line style, and its endpoints satisfy a preset connection relationship with the basic geometric primitive;

[0031] The entity object is a dotted graphic element, and the distance between its center point and the endpoint of the basic geometric element is less than a preset distance threshold.

[0032] The pointing direction of the indicator symbol entity is determined based on the main axis direction of its outer contour or the geometric features of its sharp corner vertices.

[0033] Optionally, determining the target component candidate corresponding to the target region in the component candidate set includes:

[0034] In the candidate set of components, a candidate set of pipeline components is determined based on preset pipeline element determination conditions;

[0035] When the target area and at least two pipeline component candidates in the pipeline component candidate set meet the preset overlap condition, a directional sampling band is constructed with the pointing endpoint of the visual guide primitive as the reference point and along the pointing direction of the visual guide primitive, and a transverse sampling profile sequence is obtained for the at least two pipeline component candidates within the directional sampling band.

[0036] For the at least two pipeline component candidates, boundary projection response sequences are generated based on the transverse sampling profile sequence, respectively.

[0037] A consistency score function is constructed based on the pipeline physical attribute items obtained from the source text annotation and parsing, and a consistency score is calculated for the boundary projection response sequence of each pipeline component candidate.

[0038] The pipeline component candidate with the best consistency score is selected from the at least two pipeline component candidates and used as the target component candidate corresponding to the target region.

[0039] Optionally, the construction of the consistency score function and the calculation of the consistency score include:

[0040] A set of parameters for characterizing the two-dimensional projection scale is determined from the pipeline physical property items, the set of parameters including at least one of pipe diameter parameters, cross-sectional size parameters, and elevation parameters;

[0041] A matched filter template is generated based on the parameter set. The matched filter template is used to characterize the boundary spacing and boundary polarity features of the pipeline in the transverse sampling profile sequence.

[0042] The boundary projection response sequence and the matched filter template are correlated to obtain a correlated output sequence, and the peak-to-sidelobe ratio is calculated based on the main peak-to-peak value and sidelobe statistics of the correlated output sequence.

[0043] The consistency score is obtained by weighting and fusing at least two of the main peak value, the peak-side ratio, and the main peak half-width ratio.

[0044] Optionally, establishing a strong constraint binding relationship between the source text annotation and the target component candidate includes:

[0045] For the candidate target components within the target area, obtain the boundary projection response sequence of the candidate component and the corresponding output sequence of the matched filter template.

[0046] Extract the peak value, peak position, full width at half maximum (FWHM) of the main peak, and peak-side ratio from the relevant output sequence to construct a binding feature vector;

[0047] Based on at least one of the pipe diameter parameters and cross-sectional size parameters obtained from the source text annotation parsing, the expected range of the main peak position is determined, and a generalized likelihood ratio test statistic is constructed based on the binding feature vector. The generalized likelihood ratio test statistic is used to characterize the support of the physical property consistency binding hypothesis relative to the interference binding hypothesis. The decision threshold of the generalized likelihood ratio test statistic is determined according to the preset false alarm probability.

[0048] When the generalized likelihood ratio test statistic is greater than the decision threshold and the peak position falls within the expected interval, a strong constraint binding relationship is established between the source text annotation and the target component candidate; when the generalized likelihood ratio test statistic is not greater than the decision threshold and the peak position does not fall within the expected interval, the strong constraint binding relationship is marked as unbelievable and a weak constraint candidate relationship is established for the source text annotation.

[0049] Secondly, this application provides a system for multimodal extraction and asset data generation in architectural CAD, comprising:

[0050] The acquisition module is used to acquire and parse architectural CAD files to obtain text annotation information, component geometric information, and visual guidance element information.

[0051] The first processing module is used to generate a candidate set of components based on the component geometric information; determine visual guidance primitives from the visual guidance primitive information based on the visual guidance primitive information and the text annotation information; determine the source text annotations associated with the visual guidance primitives and the target area pointed to by the visual guidance primitives; determine the target component candidates corresponding to the target area in the candidate set of components; and establish a strong constraint binding relationship between the source text annotations and the target component candidates.

[0052] The second processing module is used to establish a weak constraint candidate relationship between the source text annotation and the component candidate based on a preset association criterion for source text annotations that have not established a strong constraint binding relationship or whose strong constraint binding relationship cannot be uniquely determined.

[0053] The generation module is used to determine the candidate attribution component corresponding to the source text annotation according to the competitive arbitration rules based on the strong constraint binding relationship and the weak constraint candidate relationship, and to map the asset attribute items obtained by parsing the source text annotation to a predefined asset information structure to generate asset field values.

[0054] Compared to existing technologies, this application introduces visual guidance information that can be used to characterize directional relationships into the CAD analysis results, and constructs an attribution determination mechanism of "deterministic binding priority and candidate relationship compensation" on the component candidate set. This allows the attribution of text annotations to no longer rely solely on weak evidence such as spatial proximity, but to form more stable binding results when there is a clear direction, and to degenerate into controllable candidate relationships and complete arbitration determination when the direction is unclear or cannot be uniquely determined. This improves the robustness and consistency of automatic attribution determination under complex drawing conditions such as densely overlapping components, crowded annotations, and ambiguous directions, reduces the probability of misbinding and missed binding, improves the correctness and usability of asset field values ​​obtained from annotation analysis, and reduces the workload of manual review and correction, adapting to the large-scale processing needs of engineering asset data generation. Attached Figure Description

[0055] Figure 1 A flowchart illustrating a method for multimodal extraction and asset data generation in architectural CAD, provided as an embodiment of this application;

[0056] Figure 2 A flowchart illustrating a method for determining source text annotations and target regions, provided in an embodiment of this application;

[0057] Figure 3 A flowchart illustrating a method for determining visual guidance primitives provided in this application embodiment;

[0058] Figure 4 This is a schematic diagram of a building CAD multimodal extraction and asset data generation system provided in an embodiment of this application. Detailed Implementation

[0059] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0060] See Figure 1 The flowchart shown is a method for multimodal extraction and asset data generation in architectural CAD provided in an embodiment of this application, including steps S101 to S105, wherein:

[0061] S101: Obtain architectural CAD files and parse them to obtain text annotation information, component geometric information, and visual guidance element information;

[0062] S102: Generate a candidate set of components based on the geometric information of the components;

[0063] S103: Based on the visual guidance primitive information and the text annotation information, determine the visual guidance primitive from the visual guidance primitive information, determine the source text annotation associated with the visual guidance primitive and the target area pointed to by the visual guidance primitive, and determine the target component candidate corresponding to the target area in the component candidate set, and establish a strong constraint binding relationship between the source text annotation and the target component candidate;

[0064] S104: For source text annotations that have not established strong constraint binding relationships or whose strong constraint binding relationships cannot be uniquely determined, establish weak constraint candidate relationships between the source text annotations and the component candidates based on preset association criteria;

[0065] S105: Based on the strong constraint binding relationship and the weak constraint candidate relationship, determine the candidate belonging component corresponding to the source text annotation according to the competitive arbitration rules, and map the asset attribute items obtained by parsing the source text annotation to a predefined asset information structure to generate asset field values.

[0066] Regarding the above S101:

[0067] In practice, architectural CAD files can be vector formats such as DWG, DXF, and DGN, or intermediate representation files converted from these formats. Acquisition methods can include downloading from local storage, an enterprise document management system, or cloud storage. Before parsing, the file's encoding, unit system, and coordinate references are standardized to reduce parsing discrepancies caused by differences in drawing standards among different design institutes.

[0068] The parsing process can be executed by a CAD kernel or parsing engine. For example, the ODA / Teigha kernel, AutoCAD.NET / ObjectARX interface, or open-source DXF parsing library can be used to load the file, expand external references and block references, and apply the scaling, rotation, and translation transformations of the block references to their internal entities to obtain a unified world coordinate expression. In the case of multi-level nested blocks, the expansion can be recursively expanded to a preset upper limit of the number of levels, or the block structure can be preserved as a geometric container and the mapping relationship from the container to the entity can be output synchronously.

[0069] To improve stability, coordinates can be normalized and a geometric tolerance ε can be set to merge approximately coincident points, correct minor breakpoints, and suppress the accumulation of floating-point errors. ε can be set according to the drawing unit and typical annotation character height. For example, when the drawing unit is millimeters and the typical character height is 2.5mm to 5mm, ε can be set to 0.1mm to 0.8mm. When the drawing unit is meters and the drawing scale results in a large coordinate magnitude, ε can be converted into the corresponding length threshold according to "1% to 5% of the typical character height" so that the same logical endpoint can still be consistently identified under different file saving accuracies.

[0070] Text annotation information is used to characterize the content and geometric attributes of the annotated text. It can be extracted from text entities such as TEXT, MTEXT, and ATTRIB. Extracted fields include the text string, insertion point coordinates, rotation angle, character height, font style, layer identifier, and entity handle identifier. Furthermore, the directed bounding box or minimum bounding rectangle of the text can be calculated as the text space occupied. For text containing control characters, stacked formatting, or multi-line layout, normalization processing can be performed to output a parsable plain text sequence, such as removing format control codes, unifying full-width / half-width characters, merging line breaks, and preserving delimiters. For example, for “DN100”, Common engineering expressions such as “C35” and “H=3.600” can be generated into “attribute key-attribute value” pairs using preset regular expression patterns or lexical rules, and the correspondence between the original string and the parsing result can be recorded to maintain reversibility when traceability or rollback is needed later.

[0071] Component geometric information is used to characterize the geometric shape and topological features of building components in drawings. It can be extracted from basic geometric entities such as LINE, LWPOLYLINE, POLYLINE, ARC, CIRCLE, SPLINE, HATCH, INSERT, and their combinations. Extracted fields can include entity type, vertex sequence or parametric geometric parameters, line width / line type, closure, normal information (if it is a 3D entity), layer and handle identifier, and endpoint features, direction features, curvature features, and spatial adjacency relationships between entities can be derived. To support robust parsing of "densely overlapping" drawings, a spatial index structure (such as an R-tree or grid hash) can be established in the S101 stage to maintain approximately linear complexity when querying a given area; at the same time, the topological connection relationships between entities (such as endpoint connection, intersection, containment, tangency) can be output as the basic input for subsequent processing.

[0072] For example, when there are a large number of short line segments spliced ​​in the drawing, resulting in "virtual breaks", endpoints that are less than ε apart can be identified as connections under the tolerance of ε and recorded as the same topological node, thereby avoiding the appearance of non-real broken components in subsequent processing.

[0073] The visual guidance primitive information is used to characterize the primitives and their geometric representations in the drawing that carry semantics such as "pointing, guiding, and dimension association." It can be collected and output uniformly from two paths: explicit guidance entities and implicit guidance combinations. Explicit guidance entities can include LEADER, MLEADER, DIMENSION and their derived objects. During parsing, fields such as the vertex sequence of the guiding polyline, arrow endpoints, text attachment points, dimension line endpoints, and annotation styles can be extracted. Implicit guidance combinations can correspond to the set of geometric primitives such as leader lines / arrows / short line segments after "explosion." In stage S101, the final semantic determination may not be given directly, but the necessary candidate features and connection relationships may be output, such as candidate endpoints, candidate arrow outline parameters, candidate line type / fill style parameters, and the topological connection relationships between candidates.

[0074] In practical implementation, an endpoint connection threshold τ and a direction quantization threshold θ can be set: τ is used to determine whether two entities can be considered geometrically connected, and can be set to (1~3)×ε based on ε and the drawing scale; θ is used to suppress directional jitter caused by discretization, and can be set to 3°~12° and decrease as the drawing accuracy increases. For example, when the guide line and arrow outline have a slight deviation but still visually have the same pointing relationship, using θ=8° can keep the candidate pointing direction stable, thereby improving the consistency of subsequent processing.

[0075] Regarding S102 above:

[0076] In one embodiment, a component candidate refers to a geometric entity or combination of entities that can be identified as a potential engineering object in architectural CAD drawings. It has a computable boundary, centerline, or topological structure at the geometric level and can be assigned a unique identifier for subsequent attribution determination and attribute mapping.

[0077] The output of S102 does not require an immediate final component category or business meaning. Instead, it organizes potentially valid component geometries into a unified data structure using a pool of candidate objects, thereby avoiding accidental deletions and omissions due to overly strong prior knowledge in complex drawings.

[0078] In its implementation, S102 includes three types of processing: candidate generation, candidate merging, and candidate filtering. Candidate generation is used to form initial candidate objects from the original geometric entities; candidate merging is used to combine fragmented geometry belonging to the same actual component; and candidate filtering is used to eliminate geometric noise or auxiliary primitives that are clearly impossible to constitute an engineering component. To reduce differences caused by different drafting standards, S102 can first perform standardized preprocessing on the component geometric information, such as unifying line segments, arcs, and polylines into parametric curves or vertex sequences, and unifying style information such as curve direction, closure mark, line width, and line type into a unified field set. For geometric entities containing block references, the coordinate transformation can be completed in S101 and then directly proceed to this step to ensure that all candidate objects are under the same coordinate reference system.

[0079] The candidate generation stage can be constructed according to geometric load type. For contour-type components, closed polylines, closed-style HATCH boundaries, and closed loops formed by multiple line segments can be identified as candidate contours, and their area, perimeter, circumscribed rectangle, minimum circumscribed circle, and principal direction can be calculated. For centerline-type components, chain of line segments and polyline chains with continuous direction can be identified as candidate centerlines, and their length, curvature change, directional stability, and endpoint topological nodes can be calculated. For point-like or symbol-like components, circles, point markers, and insertion points of equipment symbol blocks that conform to a preset size range can be identified as candidate point objects, and their position and scale parameters can be output. The candidate objects of the above different load types can be represented by a unified candidate data structure, which includes at least: candidate identifier, candidate geometric representation (boundary, centerline, or point), candidate geometric feature set, source entity handle set, and candidate spatial index key value.

[0080] The candidate merging stage handles common situations where the same component is drawn separately, such as walls being spliced ​​from multiple line segments, pipelines from multiple short line segments, and equipment outlines from multiple superimposed local outline blocks. Merging can be performed based on topological connectivity and geometric consistency criteria: when two candidate objects meet the connection tolerance at their endpoints, and their directional difference is within a preset angle threshold, and their style parameters (such as layer identifier, line width, and line type) meet consistency conditions, they can be merged into the same candidate object. The merged geometry is then refitted or resampled to obtain a stable geometric representation. The connection tolerance here can follow the geometric tolerance of S101, or be adjusted by an amplification factor based on the coordinate magnitude and typical line width of the current drawing. The preset angle threshold can be set according to the expected directional continuity of the component. For example, it can be 2° to 6° for straight wall lines or straight pipeline segments, and 6° to 15% for pipelines with bends. During merging, abrupt directional changes are allowed at bends, but the abrupt change points must meet the minimum spacing constraint to avoid incorrectly connecting adjacent but unrelated line segments.

[0081] The candidate filtering stage is used to eliminate geometric objects that clearly do not constitute components, in order to control the candidate size and reduce combinatorial explosion in subsequent processing. Filtering conditions can be divided into scale filtering and structural filtering: scale filtering can eliminate objects that are too short, too small closed regions, or exceed reasonable size ranges based on parameters such as length, area, and line width; structural filtering can eliminate auxiliary lines or drawing remnants based on closure, topological isolation, and shape degradation (e.g., extremely narrow polylines that are approximately collinear, closed loops with near-zero area). Threshold settings can adopt an adaptive approach "related to the typical annotation height or typical component line width of the drawing": for example, using the text annotation height H as the scale benchmark, the minimum effective line segment length can be set to (0.5~2)×H, and the minimum effective closed region area can be set to (1~10)×H^2; when the drawing unit is millimeters and H is approximately 2.5mm~5mm, the minimum effective line segment length corresponding to the above thresholds can fall within the range of 2mm~10mm, to avoid retaining short arrow lines, fragmented dimension lines, and meaningless noise lines.

[0082] For example, for common symbol fragments in the equipment layout diagram, further filtering can be performed using an "isolation threshold": if a candidate object does not have a topological connection with other candidate objects within a preset neighborhood radius r and does not belong to the preset symbol library source, it is marked as a low-confidence candidate and can be selectively removed, where r can be set to (2~8)×H.

[0083] For example, in integrated drawings of electromechanical pipelines, the component geometry often contains a large number of parallel or overlapping centerline geometry. First, chaining line segments located within the pipeline layer set, with line widths within a preset range and continuous direction, can be merged into pipeline candidate objects. For each pipeline candidate object, the centerline direction, endpoint nodes, and local curvature distribution can be calculated. Simultaneously, the insertion points of equipment symbol blocks such as valves and tees can be used to generate point-like candidate objects, and the block name and scale parameters can be recorded as candidate features. For densely packed pipeline areas, to avoid excessive candidate merging, a "minimum separation distance" constraint can be introduced during the merging process: if two line segment chains are continuously parallel in space and the distance is less than a preset separation threshold d, but their layer identifiers or line width levels are different, cross-chain merging will not be performed. Here, d can be set to (1~3) × line width based on the line width or drawing scale.

[0084] Regarding the above S103:

[0085] In one embodiment, a visual guidance primitive can be understood as a primitive object on the drawing surface that performs the function of "pointing from a text annotation to an annotated object," possessing computable pointing end, associating end, and pointing direction characteristics; the source text annotation refers to a text object that satisfies the association rules with the visual guidance primitive in geometric connection or spatial neighborhood; the target region refers to a local spatial region determined by the pointing end of the visual guidance primitive in the drawing coordinate system, used to retrieve the pointed-to component candidate in the component candidate set. The strong constraint binding relationship output by S103 is used to characterize that, in the case of a clear pointing by the visual guidance primitive, a deterministic association pair has been established between the source text annotation and the target component candidate.

[0086] In specific implementation, S103 includes processes such as visual guidance element determination, source text annotation determination, target area determination, target component candidate determination, and binding relationship generation. First, when determining visual guidance elements from visual guidance element information, the CAD entity object set can be traversed, and a guidance candidate set can be formed by combining object type identifiers, style parameters, geometric parameters, and topological connection relationships. Entity objects that meet the preset guidance object recognition conditions are determined as visual guidance elements. Furthermore, to cover the situation where "leader lines and annotations are exploded into basic geometric element combinations" in engineering drawings, the endpoint connection features, indicator symbol connection features, text neighborhood features, and direction consistency features of basic geometric elements can be jointly determined. Geometric element combinations that meet the preset geometric topological guidance conditions are reconstructed into single visual guidance elements. Geometric normalization and attribute normalization processing are performed on visual guidance elements from different sources to ensure that they have at least unified pointing endpoints, associated endpoints, pointing directions, and element identifier fields. Secondly, when determining the source text annotation associated with the visual guidance primitive, the matching can be performed based on the bounding box neighborhood relationship between the associated endpoints of the visual guidance primitive and the text annotation object. When there is a topological connection between the visual guidance primitive and the text annotation object, the topological connection can be used as the basis for association. When there is no topological connection, the source text annotation can be determined based on spatial neighborhood matching. When there are multiple candidate text annotations, the source text annotation is selected according to the criteria of minimum distance, font scale consistency, and style consistency.

[0087] Furthermore, when determining the target area pointed to by the visual guide primitive, the pointing endpoint of the visual guide primitive can be used as a reference point, and the spatial range of the target area can be constructed by combining the pointing direction. The target area can be represented by a combination of a buffer area centered on the pointing endpoint and an directional sector along the pointing direction, so as to take into account two common drawing scenarios: the pointing endpoint falling inside the component and the pointing endpoint falling near the boundary of the component.

[0088] In the target component candidate determination stage, component candidates that have a spatial coverage relationship with the target region can be retrieved from the component candidate set, and the component candidates that satisfy the coverage relationship are selected as target component candidates. The spatial coverage relationship can be triggered by the intersection determination of the bounding box of the component candidate and the target region, and then finely verified by the inclusion relationship, nearest distance and directional consistency of the boundary curve of the component candidate and the target region, so as to reduce the risk of misselection in the case of dense components.

[0089] For example, when multiple component candidates are retrieved within the target area, comprehensive filtering can be performed based on at least two types of geometric consistency measures: first, the minimum distance from the pointing endpoint to the component candidate boundary; second, the consistency of the angle between the pointing direction of the visual guide primitive and the main direction or local normal of the component candidate; and third, the percentage of overlap between the target area and the bounding box of the component candidate. The selected target component candidates and the source text annotations together generate a strong constraint binding relationship. When the target component candidate corresponding to the strong constraint binding relationship cannot be uniquely determined under the current threshold system, the strong constraint binding relationship can be recorded as a multi-target binding state, and its corresponding target component candidate set and related geometric metric values ​​can be retained so that subsequent processing stages can continue disambiguation without changing the output format of this step.

[0090] For example, the neighborhood radius, expansion distance, connection tolerance, and angle threshold involved in S103 should be consistent with the drawing scale, text height, line width level, and coordinate magnitude to ensure robustness under different drafting units and design institute specifications. For example, the expansion distance of the text neighborhood can be set to (0.5~3)×H based on the text height H to ensure that the endpoints of the visual guide primitives can be retrieved when they fall near the text; the topology connection tolerance can be set to (0.2~1.5)×w based on the drawing coordinate accuracy and typical line width w to account for minor breakages at the endpoints and overlap deviations caused by repeated drawing; the angle threshold for directional consistency can be set to 3°~12° based on the drawing stability of the guide lines, and can be appropriately relaxed when there are a large number of oblique guide lines or isometric representations in the drawing.

[0091] For example, for a drawing object with dimensions pointing to the wall edge, the visual guidance element can be determined by the dimension line and the arrow endpoint. The source text annotation corresponds to the dimension numerical text, and the target area can be constructed by the arrow endpoint and oriented towards the wall boundary. In the component candidate set, wall outline candidates are usually represented by closed polylines or wall line bands. S103 can preferentially select the wall candidate within the target area that is closest to the arrow endpoint and whose boundary normal is consistent with the guidance direction as the target component candidate, and establish a strong constraint binding relationship between the source text annotation and the wall candidate accordingly. Through this strong constraint binding relationship, the dimension-type attribute items obtained from parsing the source text annotation can be directly assigned to the corresponding wall candidate, thereby significantly reducing the mismatch risk caused by relying solely on spatial proximity in a drawing environment with dense components and numerous annotations.

[0092] Regarding S104 above:

[0093] Weakly constrained candidate relationships are candidate association pairs between source text annotations and several component candidates. They are based on preset association criteria rather than the deterministic orientation of visually guided primitives. Therefore, these relationships usually have quantifiable association metrics to reflect the relative credibility of candidate associations.

[0094] In specific implementation, S104 can include two parts: candidate retrieval and candidate evaluation. In the candidate retrieval stage, for each source text annotation that does not form a strong constraint binding relationship, a text neighborhood retrieval region can be constructed in the map coordinate system, and candidate components that meet the spatial coverage condition with the retrieval region can be selected from the candidate component set to form an initial candidate set.

[0095] During the candidate evaluation phase, weak constraint correlation metrics can be calculated for each component candidate in the initial candidate set based on preset correlation criteria. The set of component candidates with the optimal correlation metrics is then selected as the target set of weak constraint candidate relationships. Preset correlation criteria can be composed of multiple computable constraint quantities, such as: the minimum distance from the text center point to the component candidate boundary, the overlap ratio between the text bounding box and the component bounding box, the consistency of the angle between the text baseline direction and the component's main direction, the compatibility of text style with component layer semantics, and component density penalty terms in the local region where the text is located. Among these, layer semantic compatibility can be determined through matching rules between metadata such as the component candidate's layer name, line type, color number, and professional sub-map, and the text annotation style, prefix / suffix symbols, and unit features. This ensures that weak constraint relationships maintain interpretable selection criteria even under densely overlapping conditions. To facilitate subsequent arbitration, weak constraint candidate relationships can be recorded as a set of triplets: "source text annotation identifier—component candidate identifier—correlation metric value," and the number of candidates can be limited to control the computational scale.

[0096] Regarding the setting of parameters and thresholds, the retrieval area radius of weak constraint candidate relationships, the upper limit of the number of candidates, and the screening threshold of the association metric should be related to the annotation scale of the drawing and the degree of local congestion, so as to avoid missing detection in sparse areas and introducing too many candidates in dense areas, which would lead to arbitration instability.

[0097] For example, the maximum number of candidates for each source text annotation can be set to 3 to 12, and a local density penalty can be introduced when there are too many candidates to suppress "nearby mis-adsorption in high-density areas"; the weak constraint correlation metric can be represented by a normalized score to make the metrics comparable at different drawing scales.

[0098] In the electromechanical integrated plan, when there are multiple parallel pipelines in the neighborhood of the same source text annotation, the weak constraint evaluation can prioritize the normal distance from the text center to the pipeline centerline and the consistency between the text baseline direction and the pipeline direction as the dominant constraints, and use the layer semantic compatibility as the secondary constraint, thereby obtaining several most likely pipeline component candidates.

[0099] Regarding the above S105:

[0100] In one embodiment, the competitive arbitration rule is a set of rules for prioritizing and resolving conflicts among multiple candidate attribution relationships for the same source text annotation. It at least guarantees that strong constraints are used first when they can be uniquely determined, and weak constraints are used to complete the attribution determination when strong constraints are missing or not unique, so that the attribution determination process has a definite decision path and a comparable protection scope.

[0101] In practical implementation, firstly, for each source text annotation, it is determined whether a strong constraint binding relationship exists and whether the target component candidate corresponding to this strong constraint binding relationship can be uniquely determined; if so, the target component candidate is directly determined as the assigned component candidate. Secondly, if the source text annotation does not form a strong constraint binding relationship, or if a strong constraint binding relationship corresponds to multiple different target component candidates and cannot be uniquely determined, the optimal component candidate is selected as the assigned component candidate based on the association metric value among the weak constraint candidate relationships. To improve the stability of arbitration, a relative advantage criterion for the optimal candidate can be introduced, that is, comparing the difference or ratio of the association metrics between the optimal candidate and the second-best candidate. When the relative advantage is insufficient, a secondary criterion (such as prioritizing layer semantic compatibility, directional consistency, or the smaller local density penalty) is further introduced to complete the decision, making the arbitration result insensitive to noise disturbances. Furthermore, when the same component candidate is selected by multiple source text annotations at the same time, conflict resolution can be performed based on the consistency of the field types of asset attribute items. For example, in the case of repeated assignment of the same type of field, the one with the best association metric is retained, thereby avoiding inconsistencies in field values ​​caused by multiple bindings due to local neighbors.

[0102] After determining the candidate components to be assigned, S105 can perform asset attribute parsing and structured mapping on the source text annotations. Asset attribute parsing may include: normalizing the text content, extracting numerical values ​​and units, identifying keywords and codes, and determining field categories; the predefined asset information structure can use a field dictionary to define field names, data types, and unit specifications, and map the parsed asset attribute items to structured asset field values ​​according to the field dictionary.

[0103] For example, in an electromechanical pipeline drawing, the source text annotation "DN100 1.6MPa" is parsed to obtain two asset attribute items: pipe diameter and pressure rating. After the candidate component is determined to be a pipeline component, the fields "pipe diameter = 100mm" and "pressure rating = 1.6MPa" can be generated. In a structural drawing, when the source text annotation includes material strength grade or component number, corresponding fields can be generated and written to the asset record of the corresponding component. The asset field values ​​generated above serve as structured inputs for subsequent asset data storage, verification, and linked applications, ensuring consistent field results even under complex conditions such as densely overlapping components and ambiguous annotation directions.

[0104] Optionally, in architectural CAD drawings, the connection endpoints between visually guided elements and text annotations may be affected by factors such as drawing scale, annotation density, leader line inflection points, and text bounding box overlap, resulting in unstable associations. At the same time, the target area pointed to by the visually guided elements may cover multiple component candidates, causing conflicts and making it difficult to uniquely determine strong constraint binding relationships, thereby affecting the consistency of subsequent asset field value generation.

[0105] Therefore, see Figure 2 The flowchart of a method for determining source text annotations and target regions provided in this application embodiment includes steps S201 to S202, wherein:

[0106] S201: Determine visual guidance elements from the visual guidance element information, and determine source text annotations based on the association between the visual guidance elements and the text annotation information;

[0107] S202: Determine the target area based on the pointing end of the visual guide primitive.

[0108] In this optional implementation, the parsed visual guidance primitive information may include: visual guidance primitive identifiers, geometric morphology parameters (endpoint coordinates, vertex sequence, main direction vector), style parameters, and topological connection information with other entity objects; the parsed text annotation information may include: text identifiers, insertion point coordinates, text bounding box, text height, font style, layer attributes, and parsable text content. After the processor selects the visual guidance primitives from the visual guidance primitive information, it further determines the source text annotations associated with them and the target area they point to.

[0109] Firstly, determining the source text annotation can be based on the "association relationship between visual guide elements and text annotation information." For example, the non-pointing end (also called the source end) of the visual guide element can be used as a candidate associative endpoint, and a source endpoint neighborhood can be constructed at this endpoint. The source endpoint neighborhood can be determined by the neighborhood radius centered on the source endpoint. The neighborhood radius can be adaptively determined based on the drawing scale and text height: for example, taking 1 to 3 times the text height as the baseline for the neighborhood radius, and setting upper and lower limits for very large or very small scale drawings to avoid missed associations due to an excessively small neighborhood or incorrect associations due to an excessively large neighborhood. As an example, if the text height is 2.5mm, the neighborhood radius can be 5mm to 10mm; if the CAD unit is meters and the text height is 0.0025m, the corresponding radius range can be obtained by unit conversion. Subsequently, the processor determines whether the text bounding box intersects or contains the source's neighborhood. Within the set of texts that meet the intersection or containment criteria, it calculates the association confidence score based on metrics such as "minimum distance from the endpoint to the bounding box," "distance from the source to the text insertion point," and "text layer priority." The text annotation with the highest confidence score is selected as the source text annotation. To avoid misselection due to dense text, consistency constraints can be further introduced, such as requiring the minimum distance from the source to the text bounding box to be less than a preset distance threshold, and this threshold being adjusted linearly or segmentally with the text height.

[0110] Secondly, the target area can be determined based on the pointing end of the visual guidance primitive. For example, the processor extracts the pointing endpoint of the visual guidance primitive and constructs an directional target neighborhood based on the final direction vector of the visual guidance primitive, thus giving the target area both positional and directional constraints. The directional target neighborhood can be composed of an endpoint neighborhood and a directional sector domain: the endpoint neighborhood covers the uncertainty of the landing point near the pointing end, and the directional sector domain suppresses mispointing in the opposite or lateral direction. The endpoint neighborhood radius can be set according to the line width, the length of the leader end, and the drawing scale; the half-angle of the directional sector domain can be set according to the number of leader inflection points and the drawing noise level. As an example, the endpoint neighborhood radius can be 3mm to 20mm, and the directional sector domain half-angle can be 10° to 35°; in the polyline guidance primitive, the final direction can be preferentially used as the pointing direction, and for cases where the final segment is too short, a stable direction is obtained by weighted fitting of adjacent line segments.

[0111] The target area obtained by the above method can cover the slight offset of the pointing endpoint and reduce the spread of ambiguity caused by different objects at the same point in densely populated areas.

[0112] In one alternative implementation, to address the issue that source text annotations still need to stably determine candidate components when strong constraint binding relationships are unavailable or not uniquely determined, the processor introduces a quantifiable, configurable, and deterministic arbitration process for weak constraint candidate relationships, enabling the same architectural CAD file to output consistent attribution results under different operating environments.

[0113] First, the processor establishes a candidate set data structure for the same source text annotation. Each candidate in the candidate set includes at least: a candidate component identifier, the candidate component's geometric bounding box, the coordinates of the candidate component's geometric center point, the name of the layer to which the candidate component belongs, a summary of its line type / style parameters, and a relationship type identifier when forming a candidate relationship with the source text annotation. The aforementioned geometric bounding box and geometric center point can be calculated by the CAD parsing module when traversing entity objects. For example, the minimum bounding rectangle is calculated for line segments, polylines, arcs, circles, and block reference objects based on vertex sets or geometric primitive parameters, respectively; for closed polylines, the area and principal axis direction can also be calculated additionally.

[0114] The reading of CAD entities can be achieved through ObjectARX / AutoCAD .NET API, ODA Drawings SDK, or parsing libraries such as ezdxf in DXF scenarios; geometric relationship calculations can be achieved by calling Shapely / GEOS library in Python environment, CGAL in C++ environment, and JTS in Java environment.

[0115] Secondly, when the processor determines that a strong constraint binding relationship cannot be uniquely determined, it provides executable conflict criteria and degradation triggering conditions. For example, if the number of strong constraint binding candidates corresponding to the same source text annotation is greater than one, then the processor further calculates the coverage consistency index between each candidate and the target area, for example, using the intersection ratio of the candidate component's bounding box and the target area as the coverage consistency measure. When the maximum coverage consistency is still insufficient to form uniqueness (e.g., the difference between the maximum and second largest values ​​is less than a preset difference threshold), the source text annotation is marked as a "strong constraint conflict" and enters the weak constraint arbitration process. The difference threshold can be set according to the drawing annotation density and the target area scale; the denser the drawing elements, the larger the threshold can be to suppress false uniqueness.

[0116] In the weak constraint arbitration process, the processor calculates a weak constraint score for each candidate component in the weak constraint set. The weak constraint score can be obtained by weighted fusion of three sub-scores: "geometric proximity consistency score," "semantic consistency score," and "layer / style consistency score," thereby avoiding misselection caused by relying on a single distance rule.

[0117] Firstly, the geometric proximity consistency score is used to quantify the spatial proximity between the source text annotation and the candidate component. For example, the processor calculates the minimum Euclidean distance from the bounding box of the source text annotation to the bounding box of the candidate component, and maps this distance to a score that monotonically decreases as the distance increases. This mapping can employ conventional numerical mapping methods such as exponential decay or piecewise linear decay. To ensure the mapping scale adapts to the drawing font size / scale, the distance scale parameter can be determined based on the text height. For example, a significant decay distance can be determined by taking a multiple of the text height, with a selectable range of, for example, 5 to 20 times the text height. When the CAD unit is meters or other engineering units, unit conversion can be performed before calculation.

[0118] Secondly, the semantic consistency score is used to quantify the consistency between asset attribute items obtained from source text annotation parsing and candidate component categories. For example, the processor performs rule parsing and normalization on the source text annotation content to obtain attribute item types and values. Rule parsing can be completed using regular expressions, dictionary matching, and format templates. For instance, it identifies patterns such as "DN," "φ," and "×" to obtain pipe diameter or cross-sectional dimensions, "C30 / C35" to obtain material strength grades, and "EL" and "elevation" to obtain elevation values. Then, the processor extracts category labels for candidate components. Category labels can be determined comprehensively by combining keywords from the layer name, block name, line type / color / line width combination, and geometric features (closure, aspect ratio, principal axis direction). The final semantic consistency score can be achieved using discrete scoring: a high score for a match and a low score for a non-match; or using tiered scoring: a high score for a perfect match, a medium score for a weak match, and a low score for a non-match. The correspondence between attribute item types and category labels can be permanently stored in the system as configuration data for easy auditing and maintenance.

[0119] Third, the layer / style consistency score is used to suppress misattribution across layers by leveraging prior knowledge of drafting standards. For example, the processor assigns priority weights to layers according to a preset drafting standard configuration table and sets an "allowed pairing table" to constrain the legality of combinations between source text annotation layers and candidate component layers; when a candidate component layer is not in the allowed set or the combination is invalid, the layer / style consistency score of the candidate is set to a low value or the candidate is directly eliminated, thereby improving arbitration stability.

[0120] Regarding the fusion method, the processor can weight and summarize the above sub-scores to obtain a weakly constrained score. The weight coefficients are provided by the project-level configuration and can be adaptively selected according to the drawing quality. For example, in engineering scenarios with "high layer standardization and standardized text expression," the weight of semantic and layer consistency is increased; in scenarios with "chaotic layers but more reliable spatial positions," the weight of geometric proximity consistency is increased. To avoid introducing too many parameter burdens, the specification can only provide the setting basis of "configurable weights and related to drawing standardization," and provide a set of exemplary default configurations for engineering implementation.

[0121] When the processor selects the candidate component with the highest score, to ensure deterministic output, it introduces a stable decision key for cases of "tied best or near-tied" components. For example, when the difference in weak constraint scores among multiple candidates is less than a preset tolerance threshold, the following decision keys are compared sequentially to clear the candidate: candidate component layer priority (higher priority), minimum distance from the source text label to the candidate component (smaller priority), and lexicographical or numerical order of the candidate component identifier (smaller priority). All of these decision keys are deterministic quantities that can be directly calculated or read, avoiding non-deterministic output caused by random selection or parallel uncertain sorting.

[0122] To support computational efficiency for large-scale drawings, the processor can also establish a spatial index structure for the candidate component set, enabling rapid retrieval of weakly constrained candidates. For example, an R-tree can be constructed using the bounding boxes of candidate components as index entries. When querying candidate components within the extended neighborhood of the bounding boxes of the source text annotations, candidate retrieval is reduced from a full traversal to an index retrieval. The R-tree can be implemented using libspatialindex or the rtree package in Python.

[0123] Optional, see Figure 3 The flowchart of a method for determining visual guidance primitives provided in this application embodiment includes steps S301 to S304, wherein:

[0124] S301: Traverse the set of entity objects in the architectural CAD file, and construct a guiding candidate set based on object type identifiers, style parameters, geometric parameters, and topological connection relationships between entity objects;

[0125] S302: Based on the object type identifier and style parameters, select entity objects that meet the preset guidance object identification conditions from the guidance candidate set as the first type of guidance candidate;

[0126] S303: Based on the endpoint features and topological connection features of basic geometric primitives, select a combination of geometric primitives that meets the preset geometric topological guidance conditions from the guidance candidate set as the second type of guidance candidate;

[0127] S304: Perform geometric normalization and attribute normalization on the first type of guide candidate and the second type of guide candidate to obtain visual guide primitives.

[0128] In one optional implementation, to address the issues of inconsistent data structures of visual guidance elements in architectural CAD files and the difficulty in reliably identifying guidance elements due to exploded objects or missing styles in some drawings, the system, after parsing the architectural CAD files to obtain visual guidance element information, not only directly identifies explicit guidance-type entity objects, but also performs topological inference and normalization processing on implicit guidance elements formed by the combination of basic geometric elements, thereby outputting visual guidance elements that can be consistently called by subsequent processes.

[0129] Specifically, the processor traverses the set of entity objects in the architectural CAD file. These traversed objects may include lines, polylines, arcs, splines, annotation objects, leader objects, text objects, block reference objects, and fill objects. During the traversal, the processor extracts the object type identifier, style parameters, geometric parameters, and topological connection information for each entity object, and constructs a guiding candidate set accordingly. The object type identifier characterizes the category of the entity object within the CAD's internal data structure; style parameters may include linetype, line width, color, annotation style name, text style name, block name, etc.; geometric parameters may include endpoint coordinates, vertex sequence, length, orientation angle, bounding box, curvature or radian parameters, etc.; and topological connection information characterizes the connection relationships between entity objects and other entity objects, such as endpoint coincidence, endpoints falling into the bounding box neighborhood of another object, sharing the same connection node, or group relationships formed through block references.

[0130] To reduce the impact of errors caused by different CAD formats, drawing units, and precision, the processor can introduce geometric tolerance when extracting topology connection information. For example, endpoint coincidence determination can be achieved by using a "distance between endpoints less than a preset distance threshold" method. This distance threshold can be adaptively set based on drawing units and typical element scales, such as using a certain proportion of text height, annotation arrow size, or line segment length quantile statistics in the drawing as a baseline, and then multiplying it by a preset factor to obtain the distance threshold. For example, in construction drawings using millimeters, the distance threshold can be set to 0.5mm to 3mm; in site plans using meters, the distance threshold can be set to 0.001m to 0.01m. These ranges are used to balance avoiding misclassifying independent elements as connected and avoiding omissions due to numerical precision and rounding.

[0131] After constructing the guidance candidate set, the processor selects entity objects that meet the preset guidance object identification conditions as the first type of guidance candidate based on the object type identifier and style parameters. The preset guidance object identification conditions are used to identify entity objects that have been explicitly modeled for guidance purposes at the CAD semantic level. The determination can rely on the preset guidance class set to which the type identifier belongs, and combine the annotation style parameters and line type style parameters for consistency verification. For example, when an entity object belongs to the preset type set of the guidance class or annotation class, and its style name meets the guidance purpose naming constraints defined in the project configuration file, it can be directly included in the first type of guidance candidate. When the style name is missing or has been modified non-standardly, its geometric morphological features can be further used for verification, such as whether it has a pointing end and an associated text end, whether it is accompanied by annotation text, and whether there are endpoints adjacent to the component outline, thereby improving the fault tolerance for scenarios where the type exists but the style is not standardized. In terms of implementation, DWG / DXF can obtain type and style fields through parsing libraries such as ObjectARX / AutoCAD .NET API, ODA DrawingsSDK or ezdxf; model formats such as IFC can read the component and annotation relationships through libraries such as IfcOpenShell and then convert them into a unified guide candidate data structure.

[0132] Meanwhile, based on the endpoint features and topological connectivity features of the basic geometric primitives, the processor selects combinations of geometric primitives that meet preset geometric topological guidance conditions from the guide candidate set as a second type of guide candidate. This second type of guide candidate is used to cover common engineering scenarios such as "the guide object is exploded into several basic primitives", "the guide symbol exists as a block reference or filled shape", and "the guide line and annotation text are not expressed as standard leader objects". The preset geometric topology guidance conditions can be implemented by combining three types of criteria: "endpoint adjacency," "directional consistency," and "semantic neighborhood." First, the endpoint adjacency criterion is used to confirm that a continuous path is formed between basic geometric primitives. For example, the endpoints of one line segment coincide with the endpoints of another line segment within the geometric tolerance, or the endpoints fall into the bounding box neighborhood of a symbol object. Second, the directional consistency criterion is used to confirm that the path has a clear directionality. For example, the main direction of the path remains consistent within the preset angle tolerance, avoiding the misassembly of randomly scattered short line segments into guide lines. The angle tolerance can be set according to the jitter of the drawing line type and the sampling error, and can be set to 5 degrees to 20 degrees for example. Third, the semantic neighborhood criterion is used to confirm that there is a stable association between the path and the text annotation object. For example, one end of the path is located within the extended neighborhood of the bounding box of the text annotation object. The extension distance of the extended neighborhood can be set according to the text height, and can be set to 2 to 10 times the text height for example, thus taking into account stable connections under different drawing scales.

[0133] To reduce the computational cost of combinatorial search, the processor can first create a spatial index for the basic geometric primitives and perform local clustering, then perform topological assembly and candidate generation within the clusters. For example, the bounding boxes of the basic geometric primitives can be inserted into an R-tree index, candidate adjacent edges can be obtained by querying endpoint neighborhoods, and locally connected clusters can be formed using a disjoint-set data structure or graph connectivity component algorithm. Then, within each connected cluster, the principal direction statistics and endpoint degree distribution are calculated, clusters that clearly lack directional characteristics are filtered out, and only clusters that satisfy the directional characteristics are used for the second type of guided candidate construction. Clustering and connectivity component generation can be implemented using NetworkX or a custom graph algorithm.

[0134] After obtaining the first and second types of guiding candidates, the processor performs geometric normalization and attribute normalization on both types of candidates to obtain visual guiding primitives. Geometric normalization unifies candidates from different sources into comparable geometric representations. For example, it uniformly outputs the pointing endpoints, associated endpoints, principal direction vectors, and projection lines or bounding boxes used to define the target region for the guiding primitives. When a candidate is a curve or polyline, an equivalent pointing representation can be extracted based on its endpoints and principal direction for subsequent component candidate retrieval based on the target region. Attribute normalization maps different CAD object fields to a unified set of attributes. For example, it uniformly outputs the guiding primitive identifier, source type identifier, style summary, geometric precision level, and association identifier with the text annotation object. The geometric precision level can be determined based on the candidate source; the first type of guiding candidate can be assigned a higher confidence level, while the second type of guiding candidate can be assigned a corresponding confidence level based on the number and consistency of the geometric topological criteria satisfied. To avoid excessive thresholds causing implementation burden, confidence levels can be represented by a limited number of levels, such as high, medium, and low, with the criteria for level switching provided: a higher level is obtained if more criteria are met and the directional statistics are more concentrated.

[0135] For example, in a building floor plan, the leader object is exploded into multiple short line segments with point-to-end symbols in block reference form, and some line types are edited to ordinary solid lines, making it impossible to identify the leader by relying solely on the object type identifier. In this case, the processor reconstructs a second type of leader candidate within the locally connected cluster using endpoint adjacency and direction consistency criteria, and establishes a stable association between one end of it and the text annotation object based on semantic neighborhood criteria. After geometric normalization and attribute normalization, the visual leader primitive is output, so that subsequent processes do not depend on the specific CAD object type differences when calling the visual leader primitive, thereby improving the recognition stability and consistency in complex drawings and non-standard delivered drawings.

[0136] In an optional implementation, to address the problem that the explosion of the guide primitives separates the basic geometric primitives from the indicator symbols, making it difficult to reliably identify them based solely on object type and style parameters, the system introduces preset geometric topology guidance conditions when constructing the second type of guide candidate. This involves jointly determining the connection relationships, neighborhood relationships, and directional consistency between the basic geometric primitives, indicator symbol entities, and text annotation objects. When the determination is successful, these relationships are combined and reconstructed into a single visual guide primitive, enabling subsequent processing to utilize information such as the pointing end, associated end, and main direction within a unified data structure.

[0137] Specifically, the processor first selects entity objects that meet the basic geometric primitive conditions from the guiding candidate set as objects to be determined. The basic geometric primitives may include line segments, polyline segments, sets of short line segments, etc. For each basic geometric primitive, the processor determines its endpoint sequence and determines the first end and the second end based on the endpoint sequence. The first end can be defined as "the endpoint connected to the indicator symbol entity", and the second end can be defined as "the endpoint associated with the neighborhood of the text annotation object". Before determining the connection relationship, the first end and the second end can be initially selected based on statistical characteristics such as the local neighborhood connectivity of the endpoints, the symbol density near the endpoints, and the text density near the endpoints to reduce the scale of subsequent searches.

[0138] Firstly, regarding the preset indicator symbol connection conditions: The processor establishes an endpoint neighborhood centered on the candidate endpoints of the basic geometric primitives, retrieves indicator symbol entities belonging to the preset indicator symbol set within the endpoint neighborhood, and determines whether the endpoint satisfies a connection relationship with the indicator symbol entity. This "connection relationship" can be characterized using consistent geometric adjacency criteria, such as the endpoint falling within the bounding box extended neighborhood of the indicator symbol entity, the minimum distance between the endpoint and the outer contour or key feature point of the indicator symbol entity being less than the connection tolerance, and the endpoint and the indicator symbol entity being in the same locally connected cluster. The connection tolerance can be set according to the drawing units and drawing precision: in construction drawings in millimeters, the connection tolerance can be set to 0.5mm to 2mm to tolerate explosion, snapping, and rounding errors; in general drawings in meters, the connection tolerance can be set to 0.001m to 0.01m. This tolerance setting is based on the fact that indicator symbols are usually arranged close to the leader line endpoints, and the gap between the endpoint and the symbol is much smaller than the component dimensions and annotation spacing, thus absorbing data accuracy errors without significantly introducing false detections.

[0139] Secondly, regarding the preset neighborhood conditions. The processor extracts the bounding box of the text annotation object and determines the preset neighborhood range based on the bounding box and the extension distance. This is used to determine whether the other end of the basic geometric primitive has a stable spatial relationship with the text annotation object. The bounding box can be the outer rectangle calculated from the text object's insertion point, rotation angle, text height, and string range, or it can be directly read from the bounding box field of the text object through the CAD parsing library. It is recommended that the extension distance be set based on the text height or the arrow size in the annotation style to adapt to different drawing scales: for example, the extension distance can be set to 2 to 10 times the text height; when the text density of the drawing is high and the annotations are crowded, a smaller multiple can be used to reduce false associations, and when the drawing has a long guide segment or there is a significant gap between the text and the guide line, a larger multiple can be used to reduce missed detections. When the preset neighborhood conditions are met, the second end of the basic geometric primitive should fall within the aforementioned preset neighborhood range, thereby solidifying the drafting rule that the tail end of the guide line is close to the interpreted text into a calculable geometric constraint.

[0140] Thirdly, regarding the directional consistency condition. The processor calculates the directional characteristics of the basic geometric primitives and the pointing direction of the indicator symbol entity, and determines whether they are consistent within a preset angle range. The directional characteristics of the basic geometric primitives can be determined by the direction vector from its first end to its second end; when the basic geometric primitive is a polyline or assembled from multiple line segments, the "endpoint connection direction" can be used as the main direction, and consistency verification can be performed based on the concentration of each segment's direction to avoid misidentifying non-guided primitives with obvious bends as guide lines. The pointing direction of the indicator symbol entity can be determined by the direction of its outer contour principal axis, the direction of its sharp corner vertex, or the rotation direction of the block reference. The preset angle range can be set according to the discrete error of the arrow drawing and the explosion reconstruction error in the drawing: for example, it can be set to 10 degrees to 25 degrees in common engineering drawings; it can be appropriately relaxed when the line segment is short and the directional quantification error is large, and appropriately tightened when the drawing is highly standardized and the risk of misdetection is more sensitive. The significance of this condition lies in transforming the cartographic semantics of the arrow pointing in the same direction as the guide line into an objectively measurable geometric consistency constraint, thereby suppressing mismatches where there are arrow symbols near the endpoints that are not used for guidance.

[0141] Under the conditions of preset indicator symbol connection, preset neighborhood, and direction consistency, the processor reconstructs the basic geometric primitives and the indicator symbol entities connected to them into a single visual guidance primitive. During reconstruction, geometric normalization can be performed so that the output visual guidance primitive includes at least the fields of pointing endpoint, associated endpoint, and main direction: the pointing endpoint can be the sharp corner vertex of the indicator symbol entity or the symbol feature point connected to the basic geometric primitive, the associated endpoint can be the endpoint falling into the preset neighborhood of the text annotation object, and the main direction can be determined by the pointing endpoint of the associated endpoint; at the same time, the identifiers of the combined members are aggregated and recorded so that even in scenarios where the primitive is exploded or segmented, multiple entity objects can still be uniformly expressed as a single guidance primitive that can be processed consistently downstream.

[0142] In one optional implementation, to address the problem that "different design institutes and different CAD specifications have diverse types of indicator symbols and inconsistent expression forms, making it difficult to exhaustively enumerate the indicator symbol set in advance, thus affecting the stability of implicit visual guidance primitive reconstruction", after parsing the architectural CAD file, the system adaptively filters candidate indicator symbol entities from the entity object set to form a preset indicator symbol set, and further determines the pointing direction for each indicator symbol entity that can be used for subsequent direction consistency determination.

[0143] Specifically, the processor can traverse the entity objects in the architectural CAD file, prioritize initial screening based on the object type identifier of the entity objects, and include entity objects that meet the preset indicator symbol judgment rules into the candidate pool. The preset indicator symbol judgment rules may include at least the following judgment paths, the purpose of which is to cover the two common delivery forms of "standard object expression" and "exploded basic primitive expression", and reduce missed detections caused by differences in drawing styles.

[0144] The first type of decision path targets block reference objects. After identifying an entity object as a block reference object, the processor can read its block name, insertion point, rotation angle, scale factor, and visibility status, and match the block name with a preset set of arrow block names. The preset set of arrow block names can be constructed from common engineering naming rules, such as including name patterns semantically related to "arrow, leader, dim, anno, guide, arrow", or a set of high-frequency block names obtained from historical project statistics. When the block name meets the matching conditions, the block reference object is identified as an indicator symbol entity. To improve generalization capabilities in cross-project and cross-team scenarios, block name matching can be implemented using a standardized string method, such as unifying capitalization, removing prefixes, suffixes, and separators before matching, to absorb differences caused by different naming habits.

[0145] The second type of decision path targets closed graphic objects and their fill representations. After identifying an entity as a closed graphic object, the processor can extract its fill style parameters and outer contour geometric features. Closed graphic objects can include closed polylines, closed spline curves, regions, or closed graphics with fills. Preset fill style parameters can include fill type, fill ratio, fill angle, fill pattern name, and entity color / line width, used to distinguish between "solid triangles / solid arrows used for pointing ends" and "ordinary section fills or decorative fills." Simultaneously, the processor can perform sharp corner detection on the outer contour: for example, calculating local turns in the vertex sequence of the outer contour, filtering vertices with turns significantly smaller than right angles or significantly deviating from smooth curves as sharp corner candidates, and combining this with geometric constraints such as the aspect ratio, area threshold, and number of sides of the outer contour to determine whether it is a sharp corner vertex used to represent a pointing end. The area threshold and aspect ratio threshold can be set according to the drawing unit and annotation style scale: for example, in millimeter unit drawings, the maximum side length of the bounding box of candidate indicator symbols can be limited to 0.5 to 3 times the text height to exclude large component outlines or large area filled areas from entering the candidate pool.

[0146] The third type of determination path targets short line segment elements with preset line styles. After identifying an entity as a line segment or short polyline, the processor can read its line style, line width, color, layer name, and other style parameters, and determine whether it belongs to a short line segment indicator symbol entity based on its geometric length. For example, some drawings use "slashes," "short horizontal lines," and "short dashed lines" as indicator end marks or dimension endpoint markers. These symbols are usually characterized by short length, fixed line style, and partial connection with the endpoints of guide lines. The processor can include short line segment elements with a length less than a preset length threshold and a line style belonging to a preset line style set into the candidate pool, and further determine whether they satisfy a preset connection relationship with the endpoints of basic geometric elements. The preset connection relationship can adopt a combination of endpoint distance criteria and angle criteria: for example, a specific relationship where the distance between the midpoint or endpoint of the short line segment and the endpoint of the basic geometric element is less than the connection tolerance, and the direction of the short line segment and the direction of the basic geometric element are within a preset angle range, to exclude mismatches of ordinary wall lines or grid lines.

[0147] The fourth type of determination path targets dot-marked primitives. After recognizing a polyline or dot-style marker that is a circle, an arc, or an approximate circle, the processor can extract its center point, radius, and line style, and determine whether it is used to represent an indicator end marker. To adapt to scenarios where "dots are used as leader endpoints" or "dots are used as device guide endpoints," a "distance threshold between the center point and the endpoint of the basic geometric primitive" can be used as a criterion: when the distance between the center point and the endpoint of the basic geometric primitive is less than a preset distance threshold, the dot-marked primitive is identified as an indicator symbol entity. The preset distance threshold can be determined based on the height of the annotation text or the size of the arrow: for example, it can be set to 0.2 to 1 times the text height to absorb errors caused by explosion, capture, and minor primitive offsets.

[0148] After identifying the indicator symbol entity, the processor further determines its pointing direction to support subsequent direction consistency condition determination. For block reference objects, the pointing direction can be determined by the rotation angle of the block reference object and the reference orientation in the block definition; for closed graphic objects, the pointing direction can be determined based on the direction of the outer contour principal axis and the relative position of the sharp corner vertex with respect to the centroid of the outer contour. For example, the direction from the centroid to the sharp corner vertex is taken as the pointing direction; for short line segment primitives, the pointing end orientation can be determined by taking its line segment direction and combining it with its connection position with the endpoint of the basic geometric primitive; for dot marker primitives, they can be regarded as non-directional indicator symbols, and a relaxed strategy can be adopted in subsequent direction consistency determination, such as using the direction of the basic geometric primitive as the default pointing direction, or only using the dot marker for connection conditions without participating in the direction consistency determination.

[0149] In one example scenario, a structural construction drawing defines the leader arrow as a custom block "ARW_SOLID_03," and in some drawings, the arrow is exploded into a solid triangle-filled region. The processor performs block name matching on the block reference object, and if it matches the preset arrow block name set, it is directly identified as an indicator symbol entity. For the region object, it performs fill style and sharp corner vertex detection, determining it to be a closed graphic object with solid fill and a unique sharp corner vertex, and thus includes it as a candidate indicator symbol entity. Subsequently, the processor determines the pointing direction using the block rotation angle and the centroid pointing to the sharp corner vertex, ensuring a consistent indicator direction representation across different delivery formats, thereby improving the robustness and consistency of implicit visual guidance primitive reconstruction.

[0150] In one optional implementation, to address the problem that "in architectural electromechanical CAD drawings, pipelines are densely overlapping, arranged in parallel and close proximity, and multiple pipeline component candidates fall within the same target area, making it difficult to uniquely determine the pointed pipeline based solely on spatial proximity," the system introduces a discrimination process of directional sampling zones and transverse profile sequences for pipeline components in the component candidate set. This allows the determination of target component candidates to utilize the consistency between the geometric boundary features of the pipeline and the physical attribute items of the pipeline carried in the annotation, thereby improving robustness and stability in dense pipeline layout scenarios.

[0151] Specifically, the processor can determine a candidate set of pipeline components from the candidate component set based on preset pipeline element determination conditions. These preset pipeline element determination conditions can be composed of layer semantic constraints, line type style constraints, and geometric shape constraints to avoid misclassifying non-pipeline objects such as wall lines, grid lines, and leader lines as pipelines. For example, layer semantic constraints can be established based on common electromechanical layer naming rules, such as including keyword patterns related to water supply and drainage, HVAC, fire protection, and electrical cable trays; line type style constraints can be limited to a set of styles used in projects to express pipelines, such as continuous thin lines, double-line pipeline styles, and centerline styles; geometric shape constraints can limit candidate objects to having a slenderness ratio, a fitable center direction, and approximately constant line width or double boundary spacing. For scale differences caused by different CAD drafting habits, the slenderness ratio threshold and minimum length threshold can be adaptively set in conjunction with the drawing scale, text height, or drawing frame annotation scale. For example, the minimum length threshold can be set to a multiple of the text height to exclude short, fragmented line segments.

[0152] After determining the candidate set of pipeline components, the processor can determine whether the target area and at least two pipeline component candidates in the candidate set meet a preset overlap condition to trigger subsequent fine-grained discrimination. The preset overlap condition characterizes the "competitive relationship between multiple pipeline candidates within the same target area," and can be determined by indicators such as the overlap ratio of the bounding boxes of the target area and each candidate, the number of candidate centerline segments within the target area, and the closest distance between the candidate boundary and its pointing endpoint within the target area. For example, the trigger condition can be "the number of candidates within the target area is not less than a preset number and the closest distance between at least two candidates is less than a target area scale threshold," where the target area scale threshold can be determined by the radius of the target area's circumscribed circle or the length of the short side of the smallest circumscribed rectangle of the target area's boundary.

[0153] When the preset overlap condition is met, the processor uses the pointing endpoint of the visual guide primitive as a reference point and constructs a directional sampling band along the pointing direction of the visual guide primitive. The directional sampling band can be understood as a long, thin region extending along the pointing direction, used to maintain tracking of the pointed pipeline in the "pointing direction" while simultaneously acquiring cross-sectional information that characterizes the pipeline boundary spacing and morphology in the "lateral" direction. The length and width of the directional sampling band can be set according to the drawing density and the expected pipe diameter range: the length is used to cover a stable pipeline route near the pointing endpoint, and the width is used to cover possible sets of parallel pipelines. For example, the sampling band length can be set to several times the target area scale to ensure it can overcome local occlusion or symbol interference; the sampling band width can be set to several times the expected maximum pipe diameter projection scale to cover parallel candidates and retain sufficient lateral discrimination space. The sampling step size can be set in conjunction with CAD units and line segment discrete density, for example, making the interval between adjacent cross-sections smaller than the typical text height to avoid missing pipeline offsets and local inflection points within short distances.

[0154] Within the directional sampling band, the processor acquires a sequence of lateral sampling profiles for at least two pipeline-type component candidates. The lateral sampling profiles can be generated as follows: multiple sampling positions are set along the main axis of the directional sampling band; at each sampling position, a lateral intercept line segment approximately orthogonal to the main axis is constructed; and the geometric boundaries of the candidate pipelines are sampled on this lateral intercept line segment. For pipelines expressed as bilinear lines, the geometric boundaries can be two approximately parallel boundary lines; for pipelines expressed as solid polylines, the geometric boundaries can be the outer contour boundary; for pipelines expressed as a centerline + linewidth, the equivalent boundary position can be calculated by converting the linewidth parameter with the centerline position.

[0155] After obtaining the transverse sampling profile sequence, the processor can generate a boundary projection response sequence for each pipeline component candidate based on the corresponding transverse sampling profile sequence. The boundary projection response sequence is used to normalize "multiple transverse profiles along the orientation direction" into a comparable one-dimensional sequence representation, characterizing the boundary stability and boundary spacing features of the candidate pipeline within the sampling band. For example, the processor can extract the transverse spacing of boundary pairs, the offset of the boundary pair relative to the centerline of the sampling band, and the existence marker of the boundary pairs at each transverse profile, and arrange them in sequence according to the sampling position. When some profiles are missing boundary pairs due to symbol occlusion, interpolation and confidence marking can be used to maintain consistent sequence length, while retaining missing information for subsequent scoring penalties.

[0156] Furthermore, the processor constructs a consistency scoring function based on the pipeline physical attribute items obtained from the source text annotation parsing, and calculates the consistency score for the boundary projection response sequence of each pipeline component candidate. The pipeline physical attribute items may include field information related to pipe diameter, cross-sectional dimensions, insulation thickness, elevation, and system type. Fields directly corresponding to two-dimensional projection geometry can be used to form expected constraints on boundary spacing, while fields related to system type can be used to form soft constraints on layers and line styles. To avoid introducing excessively complex mathematical expressions, the consistency scoring function can adopt a component-based evaluation approach: normalizing and fusing several evaluation components such as "the degree of matching between boundary spacing and physical attributes," "the stability of boundary spacing within the sampling band," "the degree of continuous coverage of candidate boundary pairs within the sampling band," and "the degree of consistency between candidate orientation and the pointing direction of visually guided primitives" to obtain the consistency score. The fusion weights can be adaptively set based on project experience or the reliability of annotation attributes. For example, when the text annotation explicitly gives the pipe diameter, the weight of components related to boundary spacing is increased; when the text annotation only gives the system type, the weight of components related to layer styles is increased.

[0157] Finally, the processor selects the pipeline component candidate with the best consistency score from at least two pipeline component candidates as the target component candidate corresponding to the target region. To ensure the repeatability of the output, when there are cases where the consistency scores are the same or the difference is less than a preset threshold, a stable secondary decision criterion can be introduced for deterministic clearing. For example, the candidates can be sorted by their identifier order, the effective coverage length of the candidate within the sampling band, and the projection distance along the pointing direction from the candidate's centerline to the pointing endpoint. This ensures that the same input file outputs consistent target component candidates under different operating environments.

[0158] In an example scenario, multiple parallel pipelines exist within the same corridor in the electromechanical pipeline diagram. The visual guidance element points into the overlapping area of ​​these pipelines, and the text label includes pipe diameter information such as "DN100". After constructing a directional sampling zone along the pointing direction, the system generates a transverse profile sequence for each candidate within the sampling zone, thereby obtaining the boundary projection response sequence for each candidate. Based on the pipe diameter information, the system forms an expected constraint on the boundary spacing. The system can then use the degree of matching between the boundary spacing and the expected pipe diameter as the primary scoring criterion. Simultaneously, it combines the stability of the boundary spacing within the sampling zone to suppress "local false boundaries caused by symbol interference," thus determining the target pipeline component candidate corresponding to the target area under multi-candidate competition conditions.

[0159] Optionally, to address the issue that relying solely on geometric proximity or simple boundary spacing thresholds for determination in cases of densely overlapping pipelines or inconsistent line shapes can be susceptible to local occlusion, symbol interference, and broken line segments, leading to unstable consistency among candidate pipelines, the system introduces a signal processing approach combining matched filtering and correlation analysis. This approach transforms the pipeline physical attributes obtained from source text annotation into comparable "templates," and uses indicators such as peak intensity, peak sharpness, and peak-side ratio to characterize the consistency between candidate pipelines and these templates. This results in better robustness of the consistency score against noise and local defects.

[0160] Specifically, the first step is to determine the set of parameters used to characterize the two-dimensional projection scale from the pipeline's physical properties. This parameter set may include one of the following: pipe diameter parameter, cross-sectional dimension parameter, or elevation parameter. The pipe diameter and cross-sectional dimension parameters are used to characterize the expected boundary spacing of the pipeline in the transverse sampling profile; the elevation parameter can be used to supplement the constraint on the "differences in the expression of pipelines at different elevation levels within the same corridor." For example, when the elevation information indicates that the pipeline is in an overhead or underground trench layer, the confidence level for certain line types or double-line expressions can be increased accordingly. To ensure the parameters are applicable, the parameter set can be extracted from common annotation formats by the text annotation parsing module, such as fields like "DN100," "100×50," and "±3.200," and the meaning of the dimensions is converted and normalized by combining drawing units, scale, or text height.

[0161] After obtaining the parameter set, the processor generates a matched filter template based on the parameter set. The matched filter template is used to characterize the boundary spacing and boundary polarity characteristics of the pipeline in the transverse sampling profile sequence: on the one hand, the template reflects the structure of "desired spacing between the left and right boundaries" in the transverse position; on the other hand, the template can reflect the boundary directionality, that is, the grayscale / linewidth / geometric projection responses corresponding to the two boundaries on the same profile show opposite trends at the left and right boundaries, thus forming a "polarity pair". In one implementation, the template can be constructed as a sequence with two significant responses, the interval between these two responses being determined by the pipe diameter parameter or cross-sectional size parameter; and the signs or directions of the two responses are set to opposite to match the polarity of "boundary entry / exit" or "left / right boundary" in the boundary projection response sequence. The template length can be determined according to the coverage width and sampling resolution of the transverse sampling profile, so that the template can cover the expected boundary spacing and also retain a certain boundary transition zone to suppress boundary drift caused by line segment jaggedness and discrete sampling.

[0162] Subsequently, the processor performs correlation operations on the boundary projection response sequence and the matched filter template to obtain the correlation output sequence. To reduce the impact of amplitude differences among different candidates due to line width, color, layer weight, etc., on the correlation results, the processor can normalize the boundary projection response sequence before the correlation operation, for example, by normalizing to the maximum amplitude, by energy, or by mean and variance, so that the correlation output mainly reflects the degree of shape matching rather than absolute intensity. The correlation operation can be implemented using discrete correlation or moving inner product methods, and can be performed on the CAD parsing server or the local computing terminal.

[0163] After obtaining the relevant output sequence, the processor extracts the peak value and sidelobe statistics from the sequence and calculates the peak-to-sidelobe ratio. The peak value characterizes the matching strength between the template and the candidate boundary response at the optimal alignment position; the sidelobe statistics characterize the intensity of non-target responses generated at other alignment positions besides the vicinity of the peak. To avoid including local extensions near the peak in the sidelobe count, the processor can set an exclusion interval on both sides of the peak based on the sampling resolution and the desired boundary spacing. The relevant outputs outside the exclusion interval are used as the sidelobe set, and representative statistics, such as the mean, quantiles, or maximum sidelobe value, are calculated from the sidelobe set. The peak-to-sidelobe ratio characterizes the prominence of the peak relative to the sidelobe. Its significance lies in the fact that when there are multiple similar boundary pairs in the candidate pipeline or when it is interfered with by parallel pipelines, the relevant output is prone to multiple approximate peaks, and the peak-to-sidelobe ratio decreases; when the candidate and template are highly consistent and the interference is weak, the peak is more prominent, and the peak-to-sidelobe ratio increases.

[0164] Simultaneously, the processor can extract the full width at half maximum (FWHM) of the main peak to characterize its sharpness and positioning stability. The FWHM is obtained by finding the left and right intersection points corresponding to half the peak value in the relevant output sequence and calculating the distance between them. A smaller FWHM typically indicates a clear matching position, stable boundary spacing, and a clear boundary response transition; a larger FWHM may correspond to a diffused boundary response, ambiguous pipeline representation, or the presence of superimposed interference. The calculation window for the FWHM can be combined with the lateral sampling step size setting, making this indicator comparable to changes in sampling resolution.

[0165] Finally, the processor performs a weighted fusion of at least two of the following metrics: peak value, peak-side ratio, and peak full width at half maximum (FWHM) to obtain a consistency score. Weighted fusion can be achieved using a normalized linear fusion method: for example, first mapping the peak value and peak-side ratio to a uniform scale according to the statistical range of the candidate set, and then monotonically transforming the FWHM in the direction of "the smaller the better" before participating in the fusion. The weights can be determined based on the reliability of the physical attribute items and the scene noise level: when the text annotation explicitly gives the pipe diameter or cross-sectional dimensions, the weights related to the peak value and peak-side ratio can be increased; when there are many breaks or occlusions in the drawing causing boundary response diffusion, the weight of the FWHM can be increased to suppress unstable candidates. For example, in an electromechanical pipeline diagram, if the text height is approximately 2.5 mm and the scale is 1:100, the lateral sampling step size can be set to a fraction smaller than the actual scale corresponding to the text height, so as to ensure that the half-width and peak-to-side lobe ratio calculations remain comparable under different drawing scaling conditions; the side lobe statistics can use the high quantile value outside the main peak exclusion interval to more sensitively reflect the existence of parallel interference peaks.

[0166] For example, multiple parallel water supply pipes in a corridor are represented by double lines, and the visual guidance primitives point to areas covering the boundaries of two adjacent pipes. After parsing the source text annotation to obtain "DN100", the system converts it into the expected boundary spacing under two-dimensional projection and generates a matched filter template with a double boundary structure of opposite polarity. After performing correlation operations on the boundary projection response sequences of each candidate, the correct candidate forms a single prominent main peak at the corresponding alignment position, with side lobes significantly lower than the main peak and a narrow half-width at half-maximum (HWHM); while the incorrect candidate, due to the superposition of adjacent pipe boundaries, exhibits multiple approximate peaks or diffusion of the main peak, resulting in a decreased peak-to-side-lobes ratio and an increased HWHM. Based on this, the system obtains a consistency score and achieves stable differentiation, providing a verifiable quantitative basis for subsequently selecting the optimal target component candidate among the candidates.

[0167] In one optional implementation, to address the issue that although the visual guidance primitive has pointed to the target area, there are multiple parallel or overlapping pipeline candidates within the target area, and the landing point of the guidance endpoint is affected by factors such as line width, scale, occlusion, breakage, and exploded primitives, resulting in the possibility of misbinding even if the target area is hit, the system introduces a statistical decision process based on the consistency of physical attributes when establishing a strong constraint binding relationship. On the one hand, physical attributes such as pipe diameter and cross-sectional size are used to give an interpretable expected interval constraint on the main peak position of the relevant output. On the other hand, the generalized likelihood ratio test statistic is used to determine the confidence level of "true binding" and "interference binding". Thus, unreliable bindings can be eliminated in the strong constraint binding stage, and a fallback to the weak constraint candidate relationship can be triggered.

[0168] Specifically, for candidate target components within the target area, the processor acquires the boundary projection response sequence of the candidate target component and performs correlation operations with the matched filter template to obtain the correlation output sequence. The boundary projection response sequence and the correlation output sequence can reuse the results generated from the lateral sampling profile sequence mentioned above; wherein, the independent variable of the correlation output sequence can be the lateral sampling position index, or it can be converted into lateral distance coordinates represented by the directional sampling zone reference coordinate system, which facilitates consistent threshold and interval interpretation under different drawing scaling and different sampling step sizes.

[0169] After obtaining the relevant output sequence, the processor extracts the peak value, peak position, peak half-width at half-maximum (HWHM), and peak-side ratio (PSR) from it, and constructs a bound feature vector. The peak value characterizes the maximum matching strength between the template and the candidate boundary structure; the peak position characterizes the alignment position when the maximum matching is achieved, which can be represented as the sampling index or the corresponding lateral distance; the peak HWHM characterizes the sharpness of the peak shape and the stability of its positioning; the PSR characterizes the prominence of the peak relative to the side lobes, reflecting the interference of spurious peaks caused by parallel pipelines, cross-line types, or primitive noise. To avoid the side lobe statistics being affected by the expansion near the peak, the processor can set exclusion intervals on both sides of the peak and count representative quantities of the side lobe set (e.g., high quantile values ​​or maximum side lobe values) outside the exclusion intervals, thereby obtaining the PSR. The bound feature vector can store the above indicators in a fixed order and can be accompanied by a normalization label related to the sampling resolution, making the subsequent statistical decision process transferable.

[0170] Furthermore, the processor determines the expected range of the main peak position based on at least one of the pipe diameter parameters and cross-sectional dimension parameters obtained from the source text annotation parsing. This expected range is used to constrain "where the main peak should appear when correctly bound," and its setting is based on the geometric correspondence of physical attributes in two-dimensional projection representation: for example, when the source text annotation gives the pipe diameter or cross-sectional dimension, the distance between the two boundary sides in the transverse sampling section should be consistent with the projection scale of that dimension in drawing units. The main peak position generated by the matched filter template when correctly aligned usually falls within the transverse position range that matches the projection scale. To cover common drawing and analytical errors in engineering, a tolerance band can be introduced near the theoretical scale in the expected range. The setting of the tolerance band can comprehensively consider line width, layer style, scale conversion error, sampling step size quantization error, and boundary diffusion caused by local occlusion. For example, when the drawing is a mechanical and electrical pipeline integrated plan and the text annotation can be resolved to "DN100", after the system completes the unit conversion, it can take the projection scale corresponding to "100mm" as the central scale, and introduce an extension distance on both sides of the central scale, which is determined by the line width and the sampling step size, to form the allowable interval of the main peak position; when only the cross-sectional dimensions (such as "100×50") can be resolved, the dimension consistent with the transverse sampling direction can be used as the basis for interval construction, and the other dimension can be used as an auxiliary consistency constraint for subsequent statistical decision.

[0171] Based on this, the processor constructs a generalized likelihood ratio test statistic based on the bound feature vectors. The core idea is to establish two types of hypothesis models: the physical property consistency binding hypothesis describes the typical value pattern of the bound feature vector when "the candidate is indeed the pointed pipeline and its physical properties are consistent with the text"; the interference binding hypothesis describes the value pattern of the bound feature vector when "the candidate only forms a false match due to proximity, overlap, or noise". Both types of hypothesis models can be obtained through offline calibration: for example, in a pre-collected architectural CAD sample library, a batch of "correctly bound" samples and a batch of "interference bound" samples are generated based on manual verification or reliable rules. The distribution characteristics of indicators such as peak value, half-width at half-maximum, and peak-side ratio are statistically analyzed, and these distribution characteristics are parameterized and stored as model parameters. During runtime, the processor inputs the bound feature vector of the current candidate into the two types of hypothesis models respectively, calculates its relative support under the two models, and forms the generalized likelihood ratio test statistic. The calculation of the statistics can be achieved using conventional numerical calculation procedures, such as performing likelihood estimation of each index based on Gaussian distribution, mixture distribution or histogram probability table, and fusing the support of each index in the logarithmic domain or product domain to finally obtain a single scalar statistic used for decision.

[0172] The decision threshold for the generalized likelihood ratio test statistic is determined based on a preset false alarm probability. The false alarm probability constrains the upper limit of risk associated with misclassifying interfering bindings as reliable strong bindings. Its setting can be based on the asset generation task's error tolerance requirements: when incorrect asset field values ​​lead to serious deviations in subsequent engineering quantity statistics or maintenance positioning, a lower false alarm probability can be selected to improve decision conservatism; when the task focuses more on recall, the false alarm probability can be appropriately relaxed to reduce the proportion of rollbacks to weak constraints. The threshold can be determined through offline calibration: for example, by statistically analyzing the empirical distribution of the generalized likelihood ratio test statistic on the interfering binding sample set and selecting the quantile corresponding to the preset false alarm probability as the threshold; to adapt to different drawing styles, threshold tables can also be established by grouping according to scale intervals, layer types, or sampling resolutions, and the corresponding threshold can be selected at runtime based on the current drawing metadata. For example, in MEP drawings with dense pipelines, to control the impact of misbindings on asset consistency, the false alarm probability can be set to a lower level, and the corresponding threshold can be obtained from the sample library statistics; when the drawing quality is high and visual guidance is clear, the false alarm probability can be increased to reduce unnecessary rollbacks.

[0173] Finally, a strong constraint binding relationship is established based on "two conditions": when the generalized likelihood ratio test statistic is greater than the decision threshold and the peak position falls within the expected interval, the target component candidate is determined to simultaneously satisfy "statistical confidence" and "physical scale consistency," thus establishing a strong constraint binding relationship between the source text annotation and the target component candidate. Conversely, when the above two conditions are not met, the processor marks the strong constraint binding relationship as unreliable and triggers the processing flow to establish a weak constraint candidate relationship for the source text annotation, enabling subsequent attribution determination to be re-discriminated under a more relaxed association criterion, thereby avoiding the solidification of erroneous bindings in the strong constraint stage.

[0174] In an example scenario, the pointing end of a visual guide primitive falls into the overlapping region of two parallel pipelines, both geometrically satisfying target region hit conditions. The system obtains relevant output sequences for the two candidates: one candidate has a more prominent main peak, a higher peak-to-side lobe ratio, and a smaller half-width at half-maximum. Furthermore, the main peak position matches the pipeline diameter projection scale obtained from text parsing, and the generalized likelihood ratio test statistic exceeds the threshold, thus confirming it as a credible strong binding. The other candidate also has a main peak, but its main peak position deviates from the expected interval and its side lobes are significant. Its statistic does not reach the threshold, so it is marked as unreliable and regresses to a weakly constrained candidate, thereby suppressing the risk of misbinding under overlapping conditions.

[0175] Based on the same inventive concept, this application also provides a building CAD multimodal extraction and asset data generation system corresponding to a building CAD multimodal extraction and asset data generation method. Since the principle of the system in this application is similar to the building CAD multimodal extraction and asset data generation method described above in this application, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.

[0176] Reference Figure 4 The diagram shown is a schematic of a building CAD multimodal extraction and asset data generation system provided in an embodiment of this application. The system includes:

[0177] The acquisition module 10 is used to acquire and parse architectural CAD files to obtain text annotation information, component geometric information, and visual guidance element information;

[0178] The first processing module 20 is used to generate a component candidate set based on the component geometric information; determine visual guidance primitives from the visual guidance primitive information based on the visual guidance primitive information and the text annotation information; determine the source text annotations associated with the visual guidance primitives and the target area pointed to by the visual guidance primitives; determine the target component candidates corresponding to the target area in the component candidate set; and establish a strong constraint binding relationship between the source text annotations and the target component candidates.

[0179] The second processing module 30 is used to establish a weak constraint candidate relationship between the source text annotation and the component candidate based on a preset association criterion for source text annotations that have not established a strong constraint binding relationship or whose strong constraint binding relationship cannot be uniquely determined.

[0180] The generation module 40 is used to determine the candidate attribution component corresponding to the source text annotation according to the competitive arbitration rules based on the strong constraint binding relationship and the weak constraint candidate relationship, and to map the asset attribute items obtained by parsing the source text annotation to a predefined asset information structure to generate asset field values.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for multimodal extraction and asset data generation in architectural CAD, characterized in that, The method includes: The architectural CAD file is acquired and parsed to obtain text annotation information, component geometric information, and visual guidance element information. Based on the geometric information of the components, a candidate set of components is generated; Based on the visual guidance primitive information and the text annotation information, visual guidance primitives are determined from the visual guidance primitive information, source text annotations associated with the visual guidance primitives and target areas pointed to by the visual guidance primitives are determined, and target component candidates corresponding to the target areas are determined in the component candidate set, and a strong constraint binding relationship is established between the source text annotations and the target component candidates. For source text annotations that do not have a strong constraint binding relationship or whose strong constraint binding relationship cannot be uniquely determined, a weak constraint candidate relationship is established between the source text annotation and the component candidate based on a preset association criterion. Based on the strong constraint binding relationship and the weak constraint candidate relationship, the candidate attribution component corresponding to the source text annotation is determined according to the competitive arbitration rule, and the asset attribute items obtained by parsing the source text annotation are mapped to a predefined asset information structure to generate asset field values.

2. The method for multimodal extraction and asset data generation in architectural CAD according to claim 1, characterized in that, Determining the source text annotations associated with the visual guidance primitives and the target regions pointed to by the visual guidance primitives includes: Visual guidance elements are determined from the visual guidance element information, and source text annotations are determined based on the association between the visual guidance elements and the text annotation information; the target area is determined based on the pointing end of the visual guidance elements.

3. The method for multimodal extraction and asset data generation in architectural CAD according to claim 1, characterized in that, The competitive arbitration rules include: When the strong constraint binding relationship corresponding to the source text annotation exists and is unique, the belonging component candidate is determined from the target component candidates corresponding to the strong constraint binding relationship; When the source text annotation does not establish a strong constraint binding relationship, or the strong constraint binding relationship corresponds to multiple different target component candidates and cannot be uniquely determined, the belonging component candidate is determined based on the weak constraint candidate relationship.

4. The method for multimodal extraction and asset data generation in architectural CAD according to claim 1, characterized in that, Determining the visual guidance primitive from the visual guidance primitive information includes: The entity object set of the architectural CAD file is traversed, and a guiding candidate set is constructed based on object type identifier, style parameters, geometric parameters, and topological connection relationships between entity objects; Based on the object type identifier and style parameters, entity objects that meet the preset guidance object identification conditions are selected from the guidance candidate set as the first type of guidance candidate; Based on the endpoint features and topological connection features of basic geometric primitives, combinations of geometric primitives that meet preset geometric topological guidance conditions are selected from the guidance candidate set as the second type of guidance candidate; Geometric normalization and attribute normalization are performed on the first type of guide candidate and the second type of guide candidate to obtain visual guide primitives.

5. The method for multimodal extraction and asset data generation in architectural CAD according to claim 4, characterized in that, The preset geometric topology guidance conditions include: The first end of the basic geometric primitive is connected to an indicator symbol entity in a preset set of indicator symbols, the indicator symbol entity being used to represent a pointing arrow or an indicator end mark; The second end of the basic geometric primitive is located within a preset neighborhood of the text annotation object, and the preset neighborhood is determined by the bounding box of the text annotation object and the expansion distance corresponding to the bounding box; The directional characteristics of the basic geometric primitives are consistent with the pointing direction of the indicator symbol entity within a preset angle range; Under the conditions of the preset indicator symbol connection condition, the preset neighborhood condition, and the direction consistency condition, the basic geometric primitive and the indicator symbol entities connected to it are combined and reconstructed into a single visual guidance primitive.

6. The method for multimodal extraction and asset data generation in architectural CAD according to claim 5, characterized in that, The preset set of indicator symbols is determined in the following way: From the entity objects in the architectural CAD file, select entity objects that meet the preset indicator symbol determination rules, and determine the selected entity objects as the indicator symbol entities; wherein, the preset indicator symbol determination rules include at least one of the following: The entity object is a block reference object, and its block name belongs to the preset arrow block name set; The entity object is a closed graphic object with preset fill style parameters, and its outer contour contains sharp corner vertices to represent the pointing end; The entity object is a short line segment primitive with a preset line style, and its endpoints satisfy a preset connection relationship with the basic geometric primitive; The entity object is a dotted graphic element, and the distance between its center point and the endpoint of the basic geometric element is less than a preset distance threshold. The pointing direction of the indicator symbol entity is determined based on the main axis direction of its outer contour or the geometric features of its sharp corner vertices.

7. The method for multimodal extraction and asset data generation in architectural CAD according to claim 1, characterized in that, The step of determining the target component candidate corresponding to the target region in the component candidate set includes: In the candidate set of components, a candidate set of pipeline components is determined based on preset pipeline element determination conditions; When the target area and at least two pipeline component candidates in the pipeline component candidate set meet the preset overlap condition, a directional sampling band is constructed with the pointing endpoint of the visual guide primitive as the reference point and along the pointing direction of the visual guide primitive, and a transverse sampling profile sequence is obtained for the at least two pipeline component candidates within the directional sampling band. For the at least two pipeline component candidates, boundary projection response sequences are generated based on the transverse sampling profile sequence, respectively. A consistency score function is constructed based on the pipeline physical attribute items obtained from the source text annotation and parsing, and a consistency score is calculated for the boundary projection response sequence of each pipeline component candidate. The pipeline component candidate with the best consistency score is selected from the at least two pipeline component candidates and used as the target component candidate corresponding to the target region.

8. The method for multimodal extraction and asset data generation in architectural CAD according to claim 7, characterized in that, The construction of the consistency score function and the calculation of the consistency score include: A set of parameters for characterizing the two-dimensional projection scale is determined from the pipeline physical property items, the set of parameters including at least one of pipe diameter parameters, cross-sectional size parameters, and elevation parameters; A matched filter template is generated based on the parameter set. The matched filter template is used to characterize the boundary spacing and boundary polarity features of the pipeline in the transverse sampling profile sequence. The boundary projection response sequence and the matched filter template are correlated to obtain a correlated output sequence, and the peak-to-sidelobe ratio is calculated based on the main peak-to-peak value and sidelobe statistics of the correlated output sequence. The consistency score is obtained by weighting and fusing at least two of the main peak value, the peak-side ratio, and the main peak half-width ratio.

9. The method for multimodal extraction and asset data generation in architectural CAD according to claim 8, characterized in that, The process of establishing a strong constraint binding relationship between the source text annotation and the target component candidate includes: For the candidate target components within the target area, obtain the boundary projection response sequence of the candidate component and the corresponding output sequence of the matched filter template. Extract the peak value, peak position, full width at half maximum (FWHM) of the main peak, and peak-side ratio from the relevant output sequence to construct a binding feature vector; Based on at least one of the pipe diameter parameters and cross-sectional size parameters obtained from the source text annotation parsing, the expected range of the main peak position is determined, and a generalized likelihood ratio test statistic is constructed based on the binding feature vector. The generalized likelihood ratio test statistic is used to characterize the support of the physical property consistency binding hypothesis relative to the interference binding hypothesis. The decision threshold of the generalized likelihood ratio test statistic is determined according to the preset false alarm probability. When the generalized likelihood ratio test statistic is greater than the decision threshold and the peak position falls within the expected interval, a strong constraint binding relationship is established between the source text annotation and the target component candidate; when the generalized likelihood ratio test statistic is not greater than the decision threshold and the peak position does not fall within the expected interval, the strong constraint binding relationship is marked as unbelievable and a weak constraint candidate relationship is established for the source text annotation.

10. A system for multimodal extraction and asset data generation in architectural CAD, characterized in that, include: The acquisition module is used to acquire and parse architectural CAD files to obtain text annotation information, component geometric information, and visual guidance element information. The first processing module is used to generate a candidate set of components based on the component geometric information; Based on the visual guidance primitive information and the text annotation information, visual guidance primitives are determined from the visual guidance primitive information, source text annotations associated with the visual guidance primitives and target areas pointed to by the visual guidance primitives are determined, and target component candidates corresponding to the target areas are determined in the component candidate set, and a strong constraint binding relationship is established between the source text annotations and the target component candidates. The second processing module is used to establish a weak constraint candidate relationship between the source text annotation and the component candidate based on a preset association criterion for source text annotations that have not established a strong constraint binding relationship or whose strong constraint binding relationship cannot be uniquely determined. The generation module, based on the strong constraint binding relationship and the weak constraint candidate relationship, determines the candidate attribution component corresponding to the source text annotation according to the competitive arbitration rules, and maps the asset attribute items obtained from parsing the source text annotation to a predefined asset information structure to generate asset field values.