An industrial CAD drawing multi-agent auditing method and system based on addressable layered scene graph and double-track evidence cooperation
Patent Information
- Application Number
- CN202611249415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-22
AI Technical Summary
[0007]为了解决现有工业CAD图纸审核过程中存在的结构化信息利用不足、复杂工程关系识别困难、多智能体协同机制不完善以及审核结果可追溯性和可执行性较弱的问题,本发明提供一种基于可寻址分层场景图和双轨证据协同的工业CAD图纸多智能体审核方法及系统
[0025]相较于现有工业CAD图纸审核方案在图纸信息利用不足、不同坐标来源难以统一、复杂工程语义理解不充分、多专业任务协同受限以及审核结果难以定位回写等方面的技术问题,本发明围绕CAD图纸的对象级解析、可寻址分层场景图、丰富审核邻域、确定性与推断双轨证据、多智能体受控编排及结果交付闭环形成完整处理流程,使图纸对象、工程关系、项目资料、审核任务和最终结论能够在统一标识与坐标体系下关联,在提高问题识别覆盖度的同时兼顾精确定位、证据追溯、运行控制和工程整改需求。
Smart Images

Figure CN122797342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of CAD drawing review technology, and in particular to a multi-agent review method and system for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration. Background Technology
[0002] With the development of intelligent manufacturing, digital delivery, and full lifecycle engineering management, CAD drawings are no longer just two-dimensional files for printing or viewing. Instead, they are core data carriers for conveying design intent, manufacturing constraints, and acceptance criteria between disciplines such as mechanical engineering, shipbuilding and ocean engineering, chemical processes, building structures, and electrical control. An industrial CAD drawing typically contains line segments, arcs, polylines, splines, block references, attribute definitions, dimensions, leader lines, text, tables, external references, layers, layouts, viewports, and revision history. A batch of project drawings also forms cross-file relationships through drawing numbers, equipment tag numbers, page headers, and document numbers. Any omission, incorrect numbering, broken line, version inconsistency, or specification conflict in the drawings can be amplified along the procurement, manufacturing, installation, and commissioning chain. Therefore, an industrial CAD review technology is needed that can understand the object structure, engineering semantics, and project relationships, and provide traceable evidence.
[0003] Currently, the review of industrial CAD drawings relies heavily on manual review. Engineers review drawings item by item by zooming, panning, switching layers, querying object attributes, and comparing against checklists. They can interpret complex design intentions based on engineering experience. However, review time typically increases almost linearly or even superlinearly with the number of drawings, object density, and cross-disciplinary references. When a project involves multiple version upgrades, reviewers also need to repeatedly compare previous and current versions, which can easily lead to visual fatigue, missed issues, duplicate checks, and inconsistencies in standards among different personnel.
[0004] In addition, rule-based CAD inspector auditing methods are also widely used. These methods can directly access DWG or DXF objects and are suitable for checking clearly defined aspects such as layer naming, linetypes and line weights, title block fields, tag formats, duplicate numbering, and dimension styles. Their advantages include high accuracy and fast processing. However, industry auditing rules are an open set; different drawing types, company conventions, system types, and project stages have different checkpoints, which cannot be fully expressed by a limited set of rules. Forcibly expanding the rule coverage can lead to numerous false alarms in new drawing formats due to threshold and pattern matching issues, causing engineers to lose trust in the process.
[0005] In recent years, the rapid development of intelligent agent applications has provided a new path for the review of industrial CAD drawings. However, in specific drawing review application scenarios, when a single intelligent agent simultaneously undertakes drawing parsing, connection understanding, specification retrieval, equipment list verification, cross-drawing comparison, and report generation, it is prone to problems such as confused task scopes, omissions in tool selection, mutual contamination of long contexts, and duplicate conclusions. Simply splitting the task into multiple intelligent agents can lead to problems such as different agents giving contradictory conclusions on the same object, the precise readings of deterministic rules and probabilistic opinions from model inferences being mixed together, concurrent tasks potentially exceeding computational or cost budgets, and difficulties in merging sub-task results if they lack a unified evidence structure. Existing solutions lack capability vectors, budget constraints, evidence separation, and conflict resolution mechanisms for CAD review.
[0006] Therefore, a new intelligent review solution for industrial CAD drawings is needed. While maintaining the accuracy and stable identity of the original objects, an addressable hierarchical scene graph combining spatial, topological, constraint, semantic, knowledge, and project temporal elements should be established. Rich neighborhoods with measurable coverage and adjustable dimensions should be generated according to anchor points. A language model should handle open-ended understanding, while deterministic algorithms should handle precise verification. The two types of results should be labeled as machine-determined evidence and pending inference evidence, respectively. A budget-constrained multi-agent system should complete the division of professional tasks, and the final conclusions should be reflected back to the renderer, reports, and new CAD annotation files. This solution should be able to improve the recall rate for complex problems while maintaining false alarm boundaries and engineering traceability. Summary of the Invention
[0007] To address the problems of insufficient utilization of structured information, difficulty in identifying complex engineering relationships, imperfect multi-agent collaboration mechanisms, and weak traceability and executability of review results in the existing industrial CAD drawing review process, this invention provides a multi-agent review method and system for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0008] In a first aspect, the present invention provides a multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration, which adopts the following technical solution:
[0009] A multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration includes: industrial CAD drawing access and traceable object parsing, including drawing file reception and parsing scope determination; unified modeling of native primitive objects; standardization of geometric coordinates, drawing frames and scales; standard object library and traceability mapping output;
[0010] Based on the parsing results, addressable primitives and scene graphs are constructed, including primitive node feature and spatial index construction; primitive relationship, annotation pointing and topological edge generation; engineering semantic object recognition and hierarchical aggregation; addressable primitives and scene graph output;
[0011] The review rule context generation and multi-agent task scheduling include review rule and checklist structuring; rule fit calculation and review context filtering; review task set generation and dependency construction; task agent matching and controlled scheduling.
[0012] Professional audit agent execution and dual-track evidence generation, including professional audit task context construction; professional agent audit reasoning and tool verification; deterministic inference of dual-track candidate question expression; multi-task audit result aggregation and scope verification;
[0013] Multi-source result collaborative adjudication and cross-map feedback correction, including candidate issue association grouping and duplication resolution; calculation of evidence support, conflict degree and data completeness; cross-document and cross-map cross-validation and confidence fusion; feedback correction, manual confirmation and final conclusion formation;
[0014] The process of locating, annotating, and securely delivering audit conclusions includes the structured generation of final audit conclusions; the reverse mapping of problem locations to original CAD objects; and the generation of drawing annotations, visual commands, and report content.
[0015] Secondly, an industrial CAD drawing multi-agent review system based on addressable hierarchical scene graphs and dual-track evidence collaboration includes:
[0016] The data acquisition module is configured to handle industrial CAD drawing access and traceable object parsing, including receiving drawing files and determining the parsing scope; unified modeling of native graphic primitives; standardization of geometric coordinates, drawing frames, and scales; and output of a standard object library and traceability mapping.
[0017] The primitive construction module is configured to construct addressable primitives and scene graphs based on the parsing results, including primitive node feature and spatial index construction; primitive relationship, annotation pointing and topological edge generation; engineering semantic object recognition and hierarchical aggregation; and output of addressable primitives and scene graphs.
[0018] The task scheduling module is configured to generate audit rule contexts and schedule multi-agent tasks, including structuring audit rules and checklists; calculating rule fit and filtering audit contexts; generating audit task sets and building dependencies; and matching and controlling task agents.
[0019] The evidence generation module is configured to perform professional review agent execution and dual-track evidence generation, including professional review task context construction; professional agent review reasoning and tool verification; deterministic inference of dual-track candidate question expression; and multi-task review result aggregation and scope verification.
[0020] The feedback correction module is configured to perform multi-source result collaborative adjudication and cross-map feedback correction, including candidate issue association grouping and duplication resolution; calculation of evidence support, conflict degree and data completeness; cross-document and cross-map cross-validation and confidence fusion; feedback correction, manual confirmation and final conclusion formation;
[0021] The delivery module is configured to locate, annotate, and securely deliver audit conclusions, including the structured generation of final audit conclusions; reverse mapping of problem locations to original CAD objects; and generation of drawing annotations, visual commands, and report content.
[0022] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0023] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0024] In summary, the present invention has the following beneficial technical effects:
[0025] Compared to existing industrial CAD drawing review solutions, which suffer from technical problems such as insufficient utilization of drawing information, difficulty in unifying different coordinate sources, inadequate semantic understanding of complex projects, limited collaboration among multiple professional tasks, and difficulty in locating and writing back review results, this invention forms a complete processing flow around object-level parsing of CAD drawings, addressable hierarchical scene diagrams, rich review neighborhoods, dual-track evidence of determinism and inference, controlled orchestration of multiple agents, and closed-loop delivery of results. This enables drawing objects, engineering relationships, project data, review tasks, and final conclusions to be associated under a unified identification and coordinate system, improving the coverage of problem identification while also addressing the needs of precise location, evidence traceability, operational control, and engineering rectification.
[0026] First, this invention preserves layers, block names, attributes, geometry, rotation, scale, text, title blocks, layouts, and original handles from CAD drawings during object-level parsing, and establishes a bidirectional mapping between internal identifiers and original handles. By standardizing units, scales, coordinates, title blocks, and page affiliations, rule calculations, scene diagrams, reports, and rendering use a consistent coordinate source, and the model output position is verified against the authoritative insertion point in the scene diagram. Based on this, a hierarchical scene diagram is constructed that simultaneously expresses original objects, spatial neighborhoods, pipeline topology, geometric constraints, engineering semantics, knowledge relationships, and project sequence. This enables the system to perform precise calculations of common points, paths, and constraints, analyze equipment levels, system roles, cross-page targets, and specification evidence, and trace high-level semantic conclusions back to low-level CAD entities.
[0027] Secondly, this invention assigns review items such as single diagrams, system connections, cross-diagrams, specifications, equipment information, and reports to corresponding roles through checklist capability routing, project data indexing, and task-agent matching mechanisms, and organizes execution based on task vectors, tool capabilities, scope, risks, and costs. For composite review clauses, the system automatically processes countable, searchable, and locatable objective parts on the diagram, while listing parts such as capacity, operating conditions, and materials that require external data or manual judgment. When data is missing, the system outputs input gaps instead of directly determining violations. Concurrent semaphores, maximum call levels, tokens and cost budgets, read-only batches, and access control constrain automatic operation, while the event bus organizes task dispatch, tool calls, results, and visualization actions into a verifiable execution tree.
[0028] Furthermore, this invention constructs a rich neighborhood centered on the audit anchor point, containing primitives, coordinates, text, block attributes, labeled endpoints, topology, and regional information. It adaptively expands or contracts based on coverage, object density, and truncation status, ensuring the model obtains sufficient local engineering context without excessively introducing noise. The system records tag numbers, quantities, versions, object attributes, and definitive conclusions about provable topology formation, along with inferential conclusions such as connection intent, specification interpretation, and empirical risks. It then combines object handles, bounding boxes, question types, rule clauses, source independence, and evidence support to merge and adjudicate candidate questions. Simultaneously, it utilizes semantic object differences, original content multiset differences, and project graph checks for device additions / deletions, attribute position changes, topology changes, cross-page interfaces, and version consistency, thereby reducing duplicate reports, error location, and contradictory conclusions.
[0029] Finally, this invention transforms audit conclusions into structured issue records, browser spotting and path marking, generates DOCX or PDF reports, issue-annotated DXF files, and controlled CAD copies, forming a closed-loop delivery process from discovery and location to rectification. The write-back process employs plan preview, file summary, old value verification, target whitelist, new file output, baseline snapshot, and post-write read-back invariant verification to prevent accidental modification of non-target objects; even without modification permissions, read-only location and audit reports can be delivered. Project agreements, checklists, specification retrieval, task roles, and visual styles are all configurable. New drawing formats can form draft agreements based on block names, attributes, and adjacent text, and after confirmation, migration determinism is achieved, maintaining a balance between audit accuracy, coverage, project migration, and human-machine collaboration.
[0030] Existing implementation results show that the complete method of this invention achieves a problem detection rate of 91.8%, an effective problem rate of 93.6%, a location hit rate of 98.2%, an evidence integrity rate of 96.4%, and an audit coverage rate of 94.7% across 5000 decision units, with an average audit time of 41.6 seconds per image. Compared to the scene-based single-agent method, this invention improves the problem detection rate by 7.7 percentage points and the evidence integrity rate by 7.9 percentage points; even on highly complex drawings, the problem detection rate still reaches 89.7%, and the controlled write-back test did not cause any changes to non-target key summaries. Therefore, this invention can improve the audit coverage of complex connections, specification documents, cross-image and version issues while maintaining high location accuracy and evidence integrity, and provides engineers with traceable audit conclusions. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of a multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration, according to Embodiment 1 of the present invention.
[0032] Figure 2 This is a schematic diagram comparing the problem detection rates of various methods in Embodiment 1 of the present invention;
[0033] Figure 3 This is a schematic diagram comparing the effective problem rates of various methods in Embodiment 1 of the present invention;
[0034] Figure 4 This is a schematic diagram comparing the location hit rates of various methods in Embodiment 1 of the present invention;
[0035] Figure 5 This is a schematic diagram comparing the evidence completeness rates of various methods in Embodiment 1 of the present invention;
[0036] Figure 6 This is a schematic diagram comparing the review coverage of various methods in Embodiment 1 of the present invention;
[0037] Figure 7 This is a schematic diagram comparing the average review time of each method in Embodiment 1 of the present invention;
[0038] Figure 8 This is a schematic diagram comparing the problem detection rates of various methods in Embodiment 1 of the present invention on drawings of different complexities. Detailed Implementation
[0039] The present invention will be further described in detail below with reference to the accompanying drawings.
[0040] Example 1
[0041] Reference Figure 1 This embodiment of a multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration includes:
[0042] (1) Industrial CAD drawing access and traceable object parsing module
[0043] This module receives industrial CAD drawings and their review attachments, determines the file scope, format status, and parsing boundaries for the task, and extracts primitives, layers, blocks, attributes, text, annotations, layouts, and viewport information from the original DWG or DXF object model. Unlike methods that render drawings as images and then identify them, this module saves the file summary, original handle, object owner, and drawing source at the start of parsing, enabling any nodes, relationships, and issues obtained in subsequent calculations to be traced back to the original CAD entities.
[0044] 1) Determining the scope of receiving and parsing drawing files
[0045] The system receives one or more CAD drawings submitted by a user in a single review task, and simultaneously reads the project number, drawing number, file version, encoding method, model space, layout space, and external references. For DWG files, the system preferentially calls a controlled converter to generate an intermediate DXF copy under read-only conditions; for DXF files, it directly proceeds with object traversal. The original file is always saved as an unalterable input baseline; conversion failures, version incompatibility, or file corruption are all recorded as independent statuses and are not hidden due to successful parsing of other drawings.
[0046] To clearly define the set of documents for a single audit task and their unique source, we first construct the drawing input set:
[0047] ,
[0048] in, This represents the collection of drawings pending review. Represents any drawing file in the set; Indicates the number of drawings to be entered; This represents a document summary of the corresponding drawing; This represents a function for calculating file digests. Indicates file header metadata; This indicates the function for reading the file header. The file summary serves as a common key for caching, version identification, and audit logs, enabling the differentiation of drawings with the same name but different content.
[0049] After obtaining the file set, the format validity and task input integrity rate of each file are further calculated:
[0050] ,
[0051] in, Indicates the validity status of the corresponding file; An indicator function that takes one when the condition is true; Indicates the CAD file version; This indicates the set of versions supported by the system. Indicates the file integrity verification result; Indicates file encoding; Indicates the supported encoding set; This indicates the proportion of valid inputs. The formula simultaneously checks the CAD version, file verification status, and encoding support range, and reflects the executability of the task by the proportion of all valid files to the total input files.
[0052] For files that pass the basic validation, controlled transformation and object-level parsing are performed. The parsing result is as follows:
[0053] ,
[0054] in, This represents the result of object-level resolution; Represents the CAD object parsing function; Represents a controlled format conversion function; Indicates file conversion configuration; This indicates the project's parsing configuration; This represents the set of files that failed to be parsed. This indicates an empty result. The formula applies both the transformation configuration and the project parsing configuration to the original file; files that fail to parse are added to the failure set. The system lists the reasons for failure and the unreviewed scope in the final report, and does not count failed files as passed.
[0055] After the above processing, the module outputs a drawing record with a file summary, format status, conversion source, and parsing log. This record can drive subsequent object traversal and can also be used to determine whether the cache is still valid during multiple rounds of review. When the file content, converter version, or project conventions change, the system automatically invalidates the old cache, ensuring that the review data is consistent with the current file from the input stage.
[0056] 2) Unified modeling of native primitive objects
[0057] After determining the drawing input range, the parser traverses the model space, layout space, block definitions, block references, and attribute references, reading line segments, circles, arcs, polylines, splines, fills, text, dimensions, leader lines, tables, and proxy objects. Because different CAD entities have different geometric parameters and attribute structures, this step uses object records that combine common fields with type-extended fields. This ensures that all entities have stable identities, drafting attributes, spatial ranges, and source information, while not discarding project-specific fields.
[0058] To uniformly represent different types of CAD entities, we first construct native primitive object records:
[0059] ,
[0060] in, This represents a native entity in the corresponding drawing; Indicates the system's internal stability. Represents the CAD raw handle; Indicates the entity type; Indicates the layer to which it belongs; Represents geometric parameters; Indicates the display style; Indicates text or block attributes; Represents the bounding box of an object; Indicates the object owner. Public fields are used for unified queries, while unknown fields are stored in the extended mapping.
[0061] Based on the unified object record, geometric information is broken down into continuous control points and type-specific parameters:
[0062] ,
[0063] in, A complete geometric representation of the entity; Represents the homogeneous control point matrix; Represents geometric parameters specific to the solid; Represents a homogeneous control point; Indicates the number of control points; This represents the three coordinate components of the control point. This formula allows line segments, arcs, polylines, and block references to share the coordinate transformation process, while continuing to preserve type properties such as radius, start and end angles, convexity, closure state, rotation angle, and scaling ratio through dedicated parameters.
[0064] To facilitate relational computation and semantic recognition, fixed-dimensional object input features are further constructed:
[0065] ,
[0066] in, A unified input characteristic for representing objects; Indicates the one-hot encoding function for entity types; This represents the geometric feature extraction function; Indicates the layer embedding function; This represents a style encoding function; Indicates the attribute encoding function; This represents the region location encoding function. This formula encodes object type, geometry, layer, style, attributes, and region location into uniform input features. Feature encoding is used only for computation and does not replace the original object fields; any high-level inference still retains a reference to the original handle.
[0067] After the above processing, each CAD entity forms a consistent object record with expandable content. For anonymous dynamic blocks, the system saves both the display name and the valid name; for unresolved proxy objects, at least the type, handle, owner, and display boundaries are saved, and they are explicitly marked as pending resolution so that they can still participate in spatial extent and audit coverage statistics.
[0068] 3) Standardization of geometric coordinates, drawing frames, and scales
[0069] Different industrial drawings may use millimeters, centimeters, meters, or inches, and may be drawn to actual dimensions in model space, scaled and displayed in the layout viewport, or have multiple drawing frames side by side within the same model space. Directly using the original coordinates to perform distance, connection, and cross-drawing comparisons can easily lead to scale discrepancies and page number mismatches. Therefore, this step reads the CAD header variables, viewport scale, drawing frame boundaries, and title block information to establish a unified engineering coordinate transformation for each drawing and recalculate the object space extent.
[0070] To standardize units, scales, rotations, and origin differences, a combined transformation from the original coordinates to standard engineering coordinates is constructed:
[0071] ,
[0072] in, Represents the original homogeneous coordinates; Represents standard engineering coordinates; This represents the combination transformation matrix of the corresponding drawings; Indicates translation transformation; Indicates rotational transformation; Indicates scaling transformation; Indicates the offset from the origin; Indicates the rotation correction angle; Indicates the unit conversion factor; Indicates the proportionality coefficient; This represents the standardized geometric parameters. Length-type parameters vary with scale, while angles and discrete attributes retain their semantics according to the entity type.
[0073] After the coordinate transformation is completed, the standard bounding box is recalculated based on all valid control points:
[0074] ,
[0075] in, The standard bounding box that represents an object; Represents the set of control points of an object; The x-coordinate of the standard control point; Represents the ordinate of the standard control point; This formula represents the minimum and maximum values in each coordinate direction. Instead of directly transforming the two corner points of the original bounding box, it re-calculates the minimum bounding area based on the transformed control points or curve sampling points, thus correctly handling rotated blocks, arcs, and spline curves.
[0076] For a model space containing multiple title blocks, the page assignment of objects is further calculated:
[0077] ,
[0078] in, Indicates the page to which the object belongs; This represents a framed area in the corresponding drawing; Indicates the geometric distance between the object and the drawing frame; Indicates the allowable margin at the boundary of the drawing frame; This formula indicates the drawing page assignment criteria. It prioritizes the inclusion and distance relationships between objects and the drawing frame, allowing limited allowances for objects located at the edge of the drawing frame, such as leader lines and dimension lines. Objects exceeding the entire drawing frame allowance are marked as off-page objects.
[0079] After the above processing, the original object obtains coordinates, length, and bounding box at a uniform scale, and has a clear or pending drawing page affiliation. Drawing rules, scene diagrams, problem localization, and browser rendering all use the same standard coordinates and drawing frame results, avoiding inconsistencies between the audit report page number, drawing display position, and write-back target.
[0080] 4) Standard object library and traceability mapping output
[0081] After completing object resolution and geometric standardization, the standard geometry, original attributes, drawing information, and source identity need to be reorganized into a standard object library that can be directly called by subsequent modules. This step generates an internal identifier for each object that is non-conflicting across drawings and stable when repeatedly parsed within the same file. It also establishes a bidirectional mapping between the internal identifier, file summary, original handle, drawing page, and standard bounding box, while recording the scale and resolution quality of unknown objects.
[0082] To generate stable object identities and maintain original backtracking capabilities, we first construct internal identifiers and standard objects:
[0083] ,
[0084] in, Represents the stable internal identity of an object; Represents a normalized hash function; Indicates a summary of the source file; Indicates layout, page, handle, and sub-item offset; This represents the updated standard object; This represents the function that updates the object's fields. The internal identifier is generated from the file summary, layout, page, handle, and offset of complex object sub-items. Even if different files have the same handle, there will be no conflict in the project-level object library.
[0085] A drawing object library is formed from all standard objects, and a mapping set from internal objects to original entities is established:
[0086] ,
[0087] in, A set of standard objects representing the corresponding drawings; Represents the standard objects in the collection; Indicates the number of standard objects; Represents a collection of object traceability maps; Indicates internal identity, original handle, and page to which it belongs; This represents the standard bounding box and source information. The formula represents the standard object set and its traceability mapping. Scene graph nodes, rule evidence, agent references, visualization commands, and write-back plans all reference real objects in the mapping, without relying on names temporarily generated by the model.
[0088] To determine whether the standardized results can proceed to the next stage of review, the parsing quality is further calculated:
[0089] ,
[0090] in, Indicates the quality of the drawing analysis; This indicates the weight of each quality component; Indicates the coverage ratio of object mapping; Indicates the proportion of an unknown object; This indicates the result of the frame recognition; This indicates a conditional indicator function; This indicates the validity of the coordinate transformation. The formula integrates mapping coverage, the proportion of unknown objects, the results of frame recognition, and the validity of the coordinate transformation to determine the parsing quality. If the quality is insufficient, the system can continue to provide read-only viewing, but incomplete data must not be used as the basis for definitive auditing.
[0091] After the above processing, this module outputs a standard object library, bidirectional handle mapping, layer and block definitions, title block and revision information, drawing frame area, converted and parsed versions, and a quality summary. This output constitutes the unified input for subsequent primitive-semantic hybrid representations and enables each high-level review result to return to the original CAD file along the mapping path.
[0092] Through the four steps described above, the original CAD file is transformed into a standard object library with uniform geometric scale, stable object identities, and complete source information. This process does not depend on rendering resolution, nor does it force guessing of unknown engineering objects during the parsing phase, providing a repeatable, queryable, and traceable data foundation for subsequent relationship building, semantic aggregation, and multi-agent auditing.
[0093] (2) Addressable primitives and scene graph module
[0094] This module takes the standard object library output by the previous module as input and establishes multi-layered relationships between primitive nodes, spatial indexes, annotation pointers, topological connections, geometric constraints, engineering semantics, and project knowledge. Unlike methods that only generate uninterpretable feature vectors, this module requires all high-level nodes and relationships to be saved to the paths of the constituent primitives and original handles, enabling the agent to query device roles, connection paths, and annotation affixes, as well as return actual evidence objects in the drawing.
[0095] 1) Construction of primitive node features and spatial index
[0096] The system first converts each standard object into a primitive node, with each node carrying the object type, standard geometry, layer, style, text attributes, drawing page, and source reliability. To support neighborhood queries in large-scale drawings, the system constructs a spatial index based on the standard bounding box and insertion point, and registers the drawing frame, title block, legend, equipment list, and main drawing area as region objects, allowing text with the same name to receive different semantic weights depending on its region.
[0097] To incorporate different types of objects into a unified computational space, we construct the input features of primitive nodes:
[0098] ,
[0099] in, Represents the input features of primitive nodes; Functions representing type, geometry, layer, style, attribute, and region encoding; Represents object type, layer, style, and attributes; Representing standard geometry; This represents the region where the object is located. Input features are concatenated sequentially from object type, standard geometry, layer, display style, attribute text, and region location. The original fields are still stored in the node attributes; encoded features do not replace interpretable data.
[0100] Perform non-linear encoding on the input features to obtain a set of primitive nodes carrying the real addresses:
[0101] ,
[0102] in, Represents the characteristics of primitive nodes; Represents a nonlinear activation function; Indicates node encoding parameters; Represents a set of primitive nodes; Represents an addressable primitive node; This indicates the number of nodes. The formula stores node features along with internal identifiers, original handles, and page layouts. Subsequent algorithms, even when using feature similarity, can still return the original entity via the node address.
[0103] Based on the primitive nodes, establish spatial neighborhoods for combined queries by region and object category:
[0104] ,
[0105] in, Represents the set of candidate nodes within a specified area; Indicates the query range; Represents the set of allowed object categories; Represents the complete set of primitive nodes; The standard bounding box representing a node object; This represents the node category function. The formula represents the set of nodes that intersect the query region and satisfy the category criteria. Spatial indexing first performs initial bounding box filtering, then uses true geometric distance filtering to reduce the overhead of linear scanning across the entire image.
[0106] After the above processing, the system forms a layer of graphic element nodes that can be searched by page, region, coordinate, category, and layer. The boundaries of the main drawing area, title block, legend, and equipment list are simultaneously indexed, which can reduce the error in positioning caused by confusion between equipment list text, legend symbols, and actual engineering objects.
[0107] 2) Primitive relationships, label orientation, and topological edge generation
[0108] Industrial CAD review typically relies on relationships between multiple objects, rather than individual primitive attributes. This step calculates relationships such as intersection, adjacency, parallelism, perpendicularity, collinearity, concentricity, containment, and endpoint connection based on standard geometry; for leaders and multileaders, arrowheads and text ends are extracted, and target candidates are determined by combining distance, object category, layer, and text pattern; for pipeline objects, tolerance clustering is performed on endpoints to generate topological edges with original handles.
[0109] To comprehensively describe the geometric, labeling, and regional relationships between two primitives, relationship features are constructed:
[0110] ,
[0111] in, Features representing the relationship between two primitives; Indicates the distance between objects; Indicates the difference in direction; Indicates the intersection indicator; This indicates that it contains an instruction item; Indicates layer compatibility indicators; Indicates text compatibility indicators; Indicates region compatibility. Relationship characteristics include, in order: distance, direction difference, intersection, containment, layer compatibility, text compatibility, and region compatibility. For label association, the arrowhead should be used instead of the text center.
[0112] Based on the relational features, the association strength is further calculated and the relational category is determined:
[0113] ,
[0114] in, Indicates the strength of the association between primitives; This represents the normalized activation function; Indicates the associated scoring parameters; Indicates the type of relationship; Represents the set of candidate relation categories; This represents the relationship classification parameters. The formula outputs both relationship strength and candidate categories. Categories can include spatial adjacency, label pointing, endpoint connection, geometric constraints, and compositional relationships. Low-strength relationships are not included in the deterministic review.
[0115] Only relationships that meet the reliability threshold are retained, forming a set of relationship edges with evidence sources:
[0116] ,
[0117] in, Represents the set of edges representing relationships between primitives; Represents the primitive nodes at both ends of a relational edge; Indicates the relation type; Indicates the reliability of the relationship; Evidence of a relationship; This represents the relationship retention threshold. The formula requires each relationship edge to retain the participating parties, relationship type, reliability, and source of evidence. When multiple approximate targets exist, multiple candidates are retained, and ambiguous states are handled in subsequent review stages.
[0118] After the above processing, discrete primitives are organized into a queryable relational network. Lines that visually intersect but whose endpoints do not touch will not directly form a topological connection; distant labels can establish a relationship with the target using leader arrows; each relationship can return the handle, coordinates, and measurement data of the participating primitives, providing a verifiable basis for connection review and annotation review.
[0119] 3) Engineering semantic object recognition and hierarchical aggregation
[0120] After obtaining the primitive nodes and relational edges, it is necessary to aggregate low-level objects that form the same equipment, valve, instrument, pipeline, interface, hole, title bar field, or annotation group into an engineering semantic object. This step generates candidate groups by integrating block definitions, attribute keys, tag patterns, adjacent text, object relationships, and project conventions, and then calculates the contribution of each primitive to the engineering object through attention weights. Clearly defining block attributes and manual confirmation have high reliability; model inference cannot cover the original attributes.
[0121] To determine the set of candidate primitives that constitute the same engineering object, semantic grouping is performed:
[0122] ,
[0123] in, A candidate graph tuple representing an engineering semantic object; This represents a semantic candidate recognition function; Indicates the category of the project object; Represents the set of primitive nodes and the set of relations; Represents block, text, and attribute information; This represents the semantic configuration of the project. Candidate grouping uses primitive nodes, relation edges, block information, text attributes, and project configuration simultaneously. When the recognition result is insufficient to uniquely determine the project, the system retains the candidates without forcibly generating a definitive object.
[0124] For each element in the candidate group, the contribution weight is calculated based on node characteristics and internal relationships:
[0125] ,
[0126] in, Indicates the contribution weight of elements within the candidate group; Indicates the attention parameter; Represents the characteristics of primitive nodes; This represents a summary of the relationships between primitives within a candidate group; Indicates candidate graph tuples; This indicates the exponential normalization operation. This formula gives higher contributions to primitives with clear attributes, core geometry, or key annotation relationships, while the contributions of decorative lines, repeating text, and frame elements are constrained by region and relationship.
[0127] Based on contribution weights, aggregate primitive features and internal relationships to obtain engineering semantic features:
[0128] ,
[0129] in, Represents the semantic characteristics of engineering objects; Indicates semantic aggregation parameters; Indicates the contribution weight of the primitive; Represents the characteristics of primitive nodes; Indicates the strength of relationships within the candidate group; This represents a non-linear activation function. The formula combines the internal primitive features and relational strengths of the aggregated object to output engineering semantic features suitable for auditing, while simultaneously preserving the constituent primitive set and recognition confidence.
[0130] After the above processing, the system generates engineering objects such as equipment, valves, instruments, pipelines, interfaces, title blocks, revision records, and annotation groups. Each semantic object simultaneously stores the object category, key attributes, constituent elements, identification source, and confidence level; low-confidence objects retain their original element range to avoid loss of audit scope due to semantic recognition failure.
[0131] 4) Output of addressable primitives and scene graphs
[0132] This step unifies primitive nodes, engineering semantic nodes, spatial relationships, annotation relationships, topological relationships, constraint relationships, composition relationships, and cross-graph knowledge relationships into an addressable hierarchical scene graph. Lower-level nodes provide precise geometry and raw handles, while higher-level nodes express device roles, system affiliation, parent-child relationships, specification clauses, and version relationships. Each layer is connected through stable mappings, enabling audit rules and agents to query according to the required evidence level.
[0133] To unify low-level primitives, high-level semantics, and review knowledge, a hybrid node and hybrid relation set is constructed:
[0134] ,
[0135] in, Represents a mixed set of nodes; Represents a collection of primitive, semantic, document, and rule nodes; Represents a set of mixed relations; This represents a set of primitives, components, knowledge, and version relationships. Hybrid nodes include primitives, engineering semantics, project documents, and review rules; hybrid relationships include primitive relationships, object components, engineering knowledge, and version relationships. Different layers can be built deferred as needed for queries.
[0136] Based on the hybrid nodes and relationships, a hierarchical scene graph with original object mappings is formed:
[0137] ,
[0138] in, This represents an addressable hierarchical scene diagram; Represents a mixed set of nodes; Represents a set of mixed relations; This indicates a high-level node tracing mapping; Represents the collection of underlying objects that make up higher-level nodes; This represents the object's identity, original handle, and belonging page. The mapping function in this formula expands any high-level node into the internal identifier, original handle, and belonging page of the constituent object, enabling semantic query results to directly drive problem localization.
[0139] To control the quality of the scene graph, we further calculated the node traceability rate, relationship evidence rate, and conflict situation:
[0140] ,
[0141] in, Indicates the quality of the scene graph; Indicates the weight of each quality component; Represents the traceable nodes and the entire set of nodes; Represents the relationships with evidence and the complete set of relationships; Indicates the relationship conflict rate; This represents the unknown object rate. The formula combines the proportion of traceable nodes, the proportion of relationships with evidence, the relationship conflict rate, and the proportion of unknown objects. High-level relationships of insufficient quality will not be included in the deterministic rules; they will only be provided as context for confirmation.
[0142] After the above processing, this module outputs a layered scene graph accessible through interfaces such as objects, spaces, topology, constraints, semantics, and knowledge. Subsequent modules do not need to repeatedly interpret the original DXF structure to query a device's nearby annotations, connected pipelines, system to which it belongs, cross-page targets, version relationships, and corresponding original handles.
[0143] Through the four steps described above, the standard object library is transformed into an addressable hybrid scene graph that combines low-level geometric accuracy with high-level engineering semantics. This scene graph supports both deterministic computations such as distance, endpoints, paths, and constraints, as well as intelligent understanding of device, system, and project knowledge, and all results retain the evidence chain for returning the original graph.
[0144] (3) Review rule context generation and multi-agent task scheduling module
[0145] This module transforms enterprise checklists, audit rules, user objectives, and scenario diagrams into executable audit tasks. It then filters tasks and schedules agents based on rule applicability, input completeness, expertise, and execution cost. Instead of inputting the entire diagram indiscriminately into the model, this module constructs a rich, appropriately sized context around the checklist items and audit objectives, ensuring that the task scope, required evidence, output format, and dependencies are clearly defined before execution.
[0146] 1) Structured audit rules and checklists
[0147] The system receives enterprise checklists, drafting specifications, project agreements, and user natural language objectives. It breaks down each check item into audit dimensions, applicable drawing types, target objects, necessary inputs, judgment modes, available tools, result status, and manual conditions. Items that can be directly judged through fields, quantities, dimensions, or topology are marked as deterministic capabilities; those related to connection intent, specification application, and experiential risks are marked as intelligent preliminary judgment or manual confirmation capabilities.
[0148] In order for the natural language checking items to be executed by the system, a structured rule record is first constructed:
[0149]
[0150] in, This represents a structured check item; Indicates the inspection item number, dimension, and original text; Indicates the scope of application and target audience; Indicates required input; Indicates the processing mode; Indicates the set of allowed tools; This indicates the acceptance criteria. The structured record sequentially stores the inspection item number, dimension, text, scope, target object, required inputs, capability mode, tools, and acceptance criteria. Both the original inspection text and the parsed fields are retained.
[0151] Based on the rule records, calculate the weighted completeness of the input required for this check item:
[0152] ,
[0153] in, Represents the necessary input set for the check items; This indicates a required input; Indicates the required number of inputs; Indicates the completeness of the input; Indicates the importance of the input; Indicates that the input can be an indicator function; This represents the currently available set of inputs. The formula calculates the completeness based on the importance of different inputs. When critical data such as capacity, operating conditions, materials, or manufacturer samples are missing, the system outputs insufficient data, rather than interpreting the missing information as a design error.
[0154] Based on the inspection content, available capabilities, execution risks, and costs, select the processing mode for the inspection items:
[0155] ,
[0156] in, Indicates the selected processing mode for the inspection item; Represents determinism, intelligent agents, cross-checking, and manual modes; Indicates the ability to execute patterns; Indicates the risk of pattern execution; This represents the execution cost of the model. The formula selects the model with the lowest overall cost and manageable risk among deterministic verification, agent analysis, cross-data verification, and manual processing. Deterministic facts are not left to model guessing, and open judgments are not forcibly covered by fixed rules.
[0157] Through the above processing, each inspection item has a clearly defined scope, input requirements, execution method, and result status. The system can distinguish between passed, failed, pending confirmation, insufficient data, inapplicable, and manual items, and retain the reasons for non-execution in the report, thus ensuring that the progress of the checklist truly reflects the audit coverage.
[0158] 2) Rule fit calculation and review context filtering
[0159] After obtaining the structured check items, it is necessary to determine whether they are applicable to the current map type, object type, project stage, and data scope. This step comprehensively considers map type matching, target object hit rate, necessary input completeness, tool usability, and historical reliability to calculate rule suitability. For rules that pass the screening, a rich neighborhood is constructed around the target object, tag number, user selection area, or cross-page interface, and the context scope is adjusted according to map sheet scale, path coverage, and object density.
[0160] To filter the inspection items applicable to the current audit target, the rule fit is calculated:
[0161] ,
[0162] in, Indicates the degree to which the check item fits the current drawing; Indicates the degree of matching between map types; Indicates the degree of target hit; Indicates the completeness of the input; Indicates the degree of capability availability; Indicates historical reliability; Indicates execution risk; This indicates the weight of each component. Fit is determined by map type, objective, input completeness, capability availability, historical reliability, and execution risk. Rules below the threshold are marked as inapplicable or insufficient.
[0163] For the inspection items that pass the fit screening, construct an initial neighborhood based on the audit anchor point:
[0164] ,
[0165] in, Indicates the initial audit neighborhood for the inspection item; Represents the set of audit anchor points; This indicates the standard bounding box for anchor points; Represents the minimum bounding range function; Indicates the operation of expanding the region; Indicates the map sheet scale factor; Indicates the length of the diagonal of the page to which it belongs; This indicates the upper and lower limits of the neighborhood margin. The formula centers on the area enclosed by the target object and sets the initial outer margin along the diagonal of the corresponding drawing page. Anchor points can be derived from tag numbers, object identifiers, inspection item icons, or user-selected areas.
[0166] When the number of neighboring objects exceeds the context budget, filter objects based on their relevance to the check items:
[0167] ,
[0168] in, This represents the collection of filtered review context objects; Indicates the context selection function; Indicates the candidate objects in the review neighborhood; The score represents the correlation between the object and the inspection item; Indicates the object retention threshold; This formula represents the estimated number of context objects. It prioritizes retaining target objects, connected pipelines, labeled endpoints, dimensions, tag numbers, suspected breakpoints, and related clauses, while explicitly recording the number of omitted objects to prevent the model from mistakenly assuming that objects not shown do not exist.
[0169] Through the above processing, the system creates a measurable and traceable audit context for each applicable inspection item. If path evidence extends beyond the current area, the neighborhood is expanded; if the frame, legend, or table is too noisy, the area is reduced or the anchor point is changed, thus balancing the coverage of complex relationships and the capacity of the model context.
[0170] 3) Generation of audit task sets and establishment of dependency relationships
[0171] After rule adaptation and context filtering are completed, the system generates audit tasks according to drawing-level, page-level, object-level, system-level, cross-drawing-level, and cross-document-level scopes. Each task records the target object, professional field, inspection basis, input references, risk level, expected output, and allowed tools, and calculates priority based on risk, hard error attributes, user attention, and estimated cost. Pre-tasks such as parsing, scene diagrams, data retrieval, and cross-drawing benchmark selection are explicitly connected through dependency edges.
[0172] To convert inspection items into independent executable units, construct an audit task log:
[0173] ,
[0174] in, This represents an executable audit task; Indicates the task number; Indicates the task scope and objective; Indicates the professional field and the basis of the rules; Indicates input references and risks; This indicates the permitted tools and expected outputs. The task log includes the user's identity, scope, objective, area of expertise, rules, input references, risks, available tools, and expected outputs. The task must reference real-world scenario objects or project materials.
[0175] For multiple tasks to be executed, calculate task priorities based on risk and resource requirements:
[0176] ,
[0177] in, Indicates task priority; This indicates the risk, hard error, user concern, and severity categories; Indicate the estimated cost and estimated time; This indicates the weight of each priority item. The formula prioritizes high-risk issues, hard errors, and user-critical objects, while deducting estimated costs and time. The report aggregation task is executed after the evidence task completes.
[0178] To ensure that tasks are executed in the order in which the data is generated, a task dependency graph is established:
[0179] ,
[0180] in, Represent a task dependency graph; Represents a set of task nodes; Represents the set of task-dependent edges; Indicates the conditions for starting the task; Represents the set of prerequisite tasks; Indicates the status of the preceding task; This indicates the completion and controlled skip states. The formula stipulates that a subsequent task can only begin after all preceding tasks have been completed or explicitly skipped. Failures and insufficient data do not automatically transition to a pass state; instead, the affected scope propagates along the dependency graph.
[0181] Through the above processing, single drawing parsing, equipment connection, specification retrieval, cross-drawing comparison, and result summarization are organized into a set of tasks with clear dependencies. Independent drawing or system tasks can be executed in parallel, and tasks that depend on the same object library or data index share a read-only cache, thereby reducing redundant reads.
[0182] 4) Task agent matching and controlled scheduling
[0183] After forming a set of review tasks, the system calculates the matching degree based on the domain capabilities, available tools, processing scope, permission range, current load, and historical reliability of each professional intelligent agent. The system can configure intelligent agents for drawing parsing, format review, connection review, cross-drawing review, specification retrieval, equipment data verification, and reporting, and completes scheduling under concurrency, hierarchy, token, cost, and permission restrictions.
[0184] To select a suitable specialized agent for a particular task, the task-agent matching degree is calculated:
[0185] ,
[0186] in, This indicates the degree of matching between the task and the agent; This represents the tasks to be assigned and the candidate agents; Indicates domain capability similarity; Indicates the similarity of tool sets; Indicates the scope matching degree; Indicates reliability, load, and execution risk; This indicates the weight of each matching component. The matching degree is a comprehensive measure of domain vector similarity, tool set overlap, scope matching, historical reliability, load, and risk. The system reports a capability gap when a match is insufficient.
[0187] Based on the matching degree, the global task assignment result is obtained according to the permission constraints:
[0188] ,
[0189] in, This represents the globally optimal task assignment result; Indicates the amount of task assignment instructions; This indicates the degree of matching between the task and the agent; Represents the set of permissions for an intelligent agent; Indicates task permission requirements; This indicates that the permissions satisfy the instruction function. The formula requires that each task be assigned to only one agent that meets both the permissions and input requirements, and avoids multiple high-cost tasks simultaneously occupying the same resource through global optimization.
[0190] Concurrency, hierarchy, token, and cost constraints are applied simultaneously during task execution:
[0191] ,
[0192] in, This indicates the number of tasks running at the corresponding time. Indicates the maximum number of concurrent connections; Indicates the agent invocation hierarchy; Indicates the maximum call level; This indicates the consumption and cost of the smart agent tokens; This represents the token budget and cost budget. This formula constitutes the scheduler's operating boundary. When the budget limit is reached, low-priority tasks that have not yet started are stopped, while completed results and reasons for non-execution are preserved. Unbounded recursive calls by the agent are not allowed.
[0193] Through the above processing, the system establishes an auditable scheduling mapping from tasks to specialized intelligent agents. Each dispatch, tool call, result return, failure retry, user confirmation, and cancellation operation is written to the event log; tasks involving file generation or CAD modification must be subject to access control, and intelligent agents without the necessary permissions can only generate read-only analyses or operation previews.
[0194] Through the four steps described above, the checklist and user objectives are transformed into auditing tasks with clearly defined scopes, measurable input completeness, and clear dependencies. These tasks are then executed by specialized agents with appropriate capabilities and permissions within a controlled budget. This process avoids context overload caused by undifferentiated input across the entire graph, as well as boundless concurrency and task duplication among multiple agents.
[0195] (4) Professional auditing agent execution and dual-track evidence generation module
[0196] This module provides task context to the assigned professional review agents, invokes scene graph queries, deterministic rules, document retrieval, and language models to complete the professional review, and converts the review results into candidate questions with object references and evidence states. The system records results that can be directly proven by fields, geometry, and topology separately from inference results that require model or engineer judgment, preventing probabilistic opinions from being misused as definitive facts.
[0197] 1) Professional review task context construction
[0198] Before the specialized intelligent agent begins execution, the system extracts primitives, semantic objects, spatial neighborhoods, topological paths, annotation relationships, inspection rules, and fragments of project data from the hierarchical scene graph based on the task objectives. The context is stably ordered according to page, coordinates, object type, and topological hierarchy, and explicitly identifies original facts, rule derivation, semantic inference, document references, and truncated ranges, enabling the intelligent agent to distinguish between known evidence and unprovided information.
[0199] To provide a unified input for specialized intelligent agents, task-related objects, rules, and data are first serialized:
[0200] ,
[0201] in, This indicates the context of a professional review task; Represents the object serialization function; This indicates the relationship between the object and the audit anchor point; Indicates the source of the object's data; This represents the filtered scene graph object; This represents task rules and document fragments. Each object record includes its relative relationship to the audit anchor and the data source. Rules and document fragments retain their numbers, page numbers, and document summaries. The sorting results can be used for caching and repeated runs.
[0202] When the context is insufficient or there is too much noise, update the neighborhood according to the coverage and truncation status:
[0203] ,
[0204] in, Indicates the review neighborhood for the current round; Indicates the neighborhood to be reviewed in the next round; This indicates the neighborhood margin adjustment operation; Indicates path coverage; Indicates the context truncation rate; Indicates object density; Indicates coverage and density thresholds; This represents the expansion and contraction coefficients. The formula automatically expands the region when the connection path is not covered, and shrinks the region when the object density is too high or the truncation limit is reached, preventing both narrow summaries and full-map overload.
[0205] Before model invocation, further calculate the task context quality:
[0206] ,
[0207] in, Indicates task context quality; Indicates evidence coverage; Indicates the proportion of traceable evidence; Indicates the percentage of anchor points hit; Indicates the proportion of context truncation; Indicates the proportion of irrelevant objects; This represents the weight of each quality component. The formula integrates evidence coverage, object traceability, anchor hit rate, truncation ratio, and noise ratio. When quality is insufficient, the agent requests supplementary queries instead of directly outputting a complete conclusion.
[0208] Through the above processing, each specialized agent obtains a rich context consistent with the task scope. The number of omitted objects, unresolved proxy objects, and missing data in the context are all provided explicitly, preventing the agent from interpreting input gaps as object non-existence and preserving a complete context snapshot for subsequent candidate questions.
[0209] 2) Professional intelligent agent review, reasoning, and tool verification
[0210] Specialized intelligent agents perform tasks such as format verification, connectivity analysis, specification retrieval, device data verification, cross-graph comparison, or version impact analysis based on task type. The language model is responsible for understanding open-ended design intent and organizing tool calls, while deterministic tools return the number of objects, attribute values, geometric distances, topological paths, and original clause texts. Agent output must adhere to a structured schema and reference real-world objects and document fragments within the current context.
[0211] To unify the output of different professional intelligent agents, a structured review result is constructed:
[0212] ,
[0213] in, Represents the structured output of a specialized intelligent agent; Represents the corresponding professional intelligent agent; Indicates the review task; Indicates the task context; Indicates the tools available to the intelligent agent; Indicate the claims, objectives, and candidate pathways; The structured results should include at least the issue claim, target object, candidate path, clause citation, doubts, and original confidence level. Natural language descriptions cannot replace object citations.
[0214] After the agent returns, it validates the object references, document citations, and output patterns.
[0215] ,
[0216] in, This indicates the reference verification rate output by the intelligent agent; Indicates the effective indicator quantity of the structural pattern; This indicates the collection of objects to be referenced in the output; Represents the set of real nodes in the scene graph; This indicates the output citation set; This represents the set of real data retrieved by the task. The formula calculates the true hit rate for scene graph object references and project data citations, respectively, and requires the output to conform to a predefined field pattern. Non-existent handles and unretrieved terms are discarded.
[0217] Adjust the model confidence based on the citation verification results, and determine whether the candidate results can enter the evidence stage:
[0218] ,
[0219] in, Indicates the confidence level after verification; This represents the original confidence level of the model; Indicates the citation verification rate; Indicates the conditions for accepting results from the agent; Indicates the minimum verification threshold; Indicates the target object of the candidate results; This indicates an empty target. The formula uses the validation rate to adjust the model confidence. Results with an empty target, insufficient reference hits, or failed structural repairs are retained as run records but are not included in the formal candidate problem set.
[0220] After the above processing, the reasoning results of the professional intelligent agent are constrained within the scope of the real scene diagram and project data. If the coordinates given by the model correspond to the actual position number, the system uses the authoritative insertion point in the object library for reconciliation; speculative intersections that cannot be verified are retained as doubtful points and marked as pending confirmation, and are not written as confirmed objects.
[0221] 3) Deterministic Inference Dual-Track Candidate Problem Expression
[0222] To differentiate between different strengths of proof, this step establishes a deterministic evidence track and an inferential evidence track. Results such as identification number, quantity, version, object attributes, dimensional residuals, and provable topology enter the deterministic track when the input is complete and the decision boundaries are clear. Connection intent, normative interpretation, empirical risks, and visual doubts enter the inferential track, and are required to be manually verified. Both types of evidence share the target object and location fields but have different source levels and delivery permissions.
[0223] For conclusions that can be directly calculated and retested, construct a record of definitive evidence:
[0224] ,
[0225] in, Records indicating definitive evidence; Indicate the conclusion, claim, and target audience; Indicates the execution rule; Represents the object handle and bounding box; This indicates the measurement residual and the allowable threshold; Indicates the basis of the terms; This indicates a manually verified signature. A deterministic record includes rules, object handles, bounding boxes, measurement residuals, allowable thresholds, and clause sources. Deterministic conclusions cannot be generated if input is missing.
[0226] For conclusions that require semantic understanding or engineering judgment, construct a record of inferential evidence:
[0227] ,
[0228] in, Indicates inference of evidence records; Indicates the inference, claim, and target; Represents the inference model; Indicates a context snapshot; Represents the actual collection of references; Indicates the confidence level of the inference; Indicates the uncertainty of the inference; This indicates a manual confirmation flag. The inference record includes the model, context snapshot, true references, confidence level, and uncertainty, and is permanently written to the pending confirmation state by the system; the model cannot cancel this status itself.
[0229] Based on the two types of evidence, a unified record of candidate questions is generated:
[0230] ,
[0231] in, This represents a candidate question record; Indicates the candidate question number; State the problem, claims, and objectives; Indicates the level of evidence and its processing status; Indicates the confidence level of the question; Represents a set of evidence; Indicates location information; The document outlines rectification suggestions. Candidate issues are uniformly formatted, including claims, objectives, evidence level, status, confidence level, evidence, location, and rectification suggestions. This standardized structure facilitates subsequent merging but does not eliminate discrepancies in evidence tracks.
[0232] After the above processing, model opinions and machine measurements are no longer mixed in the same list of indiscriminate questions. The front-end, reporting, and write-back modules use different colors, wording, and permissions according to the evidence level: deterministic questions can be directly retested, inference questions require manual confirmation, and insufficient data items only prompt for supplementary input. Unconfirmed inferences must not drive CAD modifications.
[0233] 4) Summary of multi-task audit results and scope verification
[0234] After multiple specialized agents complete their tasks, the system summarizes candidate issues by drawing, sheet, check item, system, and cross-drawing scope, while retaining task status, reasons for non-execution, and runtime logs. The summarization process only merges structured artifact references and does not regenerate audit facts; for failed, canceled, budget-depleted, or data-insufficient tasks, the affected check items and drawing scopes are calculated and listed separately in the summarization results.
[0235] To aggregate candidate problems from different agents and tasks, a full candidate set and a task state set are constructed:
[0236] ,
[0237] in, Represents the set of all candidate problems for intelligent agents; This represents the set of intelligent agents participating in the review; This represents the set of tasks performed by the corresponding intelligent agent; This represents the candidate questions generated by the corresponding task; Represents a set of task states; Indicates the task execution status; Indicates the number of tasks. The full candidate set consists of structured questions output by each agent within the scope of each task, and the aggregator must not add new claims outside of these records.
[0238] When aggregating candidate results, the review coverage rate is calculated and a set of incomplete tasks is formed:
[0239] ,
[0240] in, Indicates the coverage rate of the audit tasks; Indicates the importance of the inspection task; Indicates a function indicating completion or product validity; Indicates the task status; Indicates the task output; This represents the set of incomplete tasks; Indicates an incomplete review task; This indicates the task completion status. The formula calculates the coverage ratio of completed and valid deliverables based on the importance of the check items. Failed, canceled, and insufficient data tasks are added to the gap set and associated with the affected drawings and rules.
[0241] To ensure the results can be verified, a task execution audit tree will be further established:
[0242] ,
[0243] in, This represents the task execution audit tree; Indicates the current running and parent running identifiers; Indicate the event type and tool; Indicates the input summary; Indicates a reference to the output artifact; Indicates the time of the event; This indicates execution permissions. The formula stores the run, parent run, event, tool, input summary, output reference, time, and permissions. Any candidate issue can be traced back to the task that generated it, tool calls, and context snapshot.
[0244] After the above processing, the system outputs a summary result including the scope of review, task status, candidate issues, and operational evidence. This result can clearly distinguish between four states: system confirmed as normal, no issues found, system not yet executed, and insufficient data for judgment. This provides complete input for the next module's candidate grouping, conflict resolution, and manual review.
[0245] Through the four steps described above, the specialized intelligent agent completes the review process under rich context and deterministic tool constraints, and outputs candidate questions that are locatable, verifiable in origin, and distinguishable in state. This dual-track evidence structure allows high-precision computation and open model understanding to each undertake their appropriate tasks, while preventing uncertain inferences from exceeding the boundaries of engineering evidence.
[0246] (5) Multi-source result collaborative adjudication and cross-graph feedback correction module
[0247] This module handles candidate issues arising from different agents, rule tools, project data, and cross-graph comparisons. It arrives at a final review conclusion through object-level grouping, evidence support and conflict measurement, multiple rounds of cross-validation, and feedback correction. The adjudication process retains all original candidate and evidence events. Duplicate expressions from the same source are discounted for relevance. Mutually exclusive claims are prioritized for processing using verifiable tools; if these methods fail, the case is transferred to manual verification.
[0248] 1) Candidate problem association grouping and repeated resolution
[0249] Different agents may describe the same problem from the perspectives of rules, connections, specifications, or cross-graph approaches, and may also make claims that are textually similar but actually different on adjacent objects. This step integrates the original handle set, standard bounding boxes, problem types, rule clauses, check items, and claim semantics to calculate the degree of association among candidates, and then forms problem groups based on the connected components of the association graph. Candidates that are only textually similar but differ in objects are not directly merged.
[0250] To determine whether two candidate questions describe the same engineering issue, the candidate correlation degree is calculated:
[0251] ,
[0252] in, Indicates the degree of correlation between two candidate questions; Indicates the similarity of sets of object handles; Indicates the overlap of the position boxes; Indicates a consistent type or rule; Indicate the semantic similarity of the claims; This indicates the weight of each related component. The degree of association is a comprehensive measure of handle set similarity, bounding box overlap, question type, rule consistency, and claim semantics. Object evidence has a higher weight than plain text similarity.
[0253] Establish a candidate question association diagram based on the degree of relevance and the scope of review:
[0254] ,
[0255] in, Represents the relationship graph of candidate problems; Represents the set of candidate problem nodes; Represents the set of candidate associated edges; Indicates the degree of correlation between candidates; Indicates the threshold for retaining associations; This indicates the scope of review for two candidates. The formula only connects candidates that meet the association threshold and have the same scope. Cross-graph and single-graph issues are handled separately according to their respective target object scopes, even if the text is similar.
[0256] Identify the connected components of the association graph, form candidate problem groups, and select representative records:
[0257] ,
[0258] in, Represent a group of candidate questions; Represents the connected component functions of an association graph; This indicates the record of the candidate group representative; Indicates the quality of candidate evidence; Indicates the completeness of candidate information; This indicates the candidate duplication penalty. The formula selects representative records based on evidence quality, information completeness, and duplication penalty, while the remaining candidates are retained as supporting or opposing sources, and the original event is not deleted due to deduplication.
[0259] Through the above processing, duplicate statements are grouped into the same candidate group, while issues with different locations, objects, or scopes remain independent. The final number of issues is calculated by candidate group, but issue details can be expanded to view all sources, timelines, and evidence objects, thus balancing report conciseness with audit completeness.
[0260] 2) Calculation of Evidence Support, Conflict, and Data Completeness
[0261] After candidate grouping is completed, it is necessary to evaluate the degree of support, opposition, and data completeness of each source within the group. This step calculates the support weight based on evidence level, source reliability, source independence, object hit rate, and citation completeness; calculates the conflict degree for mutually exclusive categories, opposite conclusions, numerical out-of-limit and normal relationships; and calculates the input gap for groups lacking drawings, clauses, or external parameters.
[0262] To quantify the evidence supporting and opposing the current claim in the candidate group, support weights and opposition weights are calculated separately:
[0263] ,
[0264] in, Indicates the support weight for the candidate group; Indicates the weight of the candidate group's opposition; Indicates the candidate question group; Indicates the independence of the source of evidence; Indicates the weight of the reliability of the evidence; Indicates the quality of evidence; This indicates a support or mutual exclusion relationship. Evidence weights consider source hierarchy, historical reliability, citation quality, and source independence simultaneously. Multiple restates of the same context within the same model use lower independence weights.
[0265] Based on the support and opposition weights, the degree of conflict and the degree of source independence are further calculated:
[0266] ,
[0267] in, Indicates the degree of conflict among candidate groups; Indicates support weights and opposition weights; Indicates the degree of independence of the source; Indicate the degree of relevance between two sources of evidence; Indicates the size of the candidate group; This represents a small constant to prevent the denominator from being zero. The formula expresses conflict as a proportion of the total weight to the opposing weight and assesses the independence of evidence using the relevance of the sources. When conflict is high, the conclusion from the text with higher confidence cannot be directly selected.
[0268] To determine whether a candidate meets the delivery requirements, the completeness of evidence regarding the object, basis, and status is calculated:
[0269] ,
[0270] in, Indicates the completeness of evidence for candidate questions; Indicates the complete indication of the target object; Indicates the complete pointer to the original handle; Indicates a complete positioning indication quantity; The rules are based on the complete instruction quantity; Indicates the complete quantity of the clause citation; This indicates the completeness of the verification status. The formula checks the target object, original handle, location, rule or clause, citation, and verification status. If the completeness is insufficient, the system supplements the verification task or fills the gaps in the output data.
[0271] After the above processing, each candidate group obtains support, conflict, independence, and completeness. Deterministic measurements, independent project data, and human verification have high evidentiary weight; repeated model calls do not artificially inflate support; normative judgments lacking necessary parameters maintain a state of insufficient data, rather than participating in a binary competition of pass or fail.
[0272] 3) Cross-document and cross-drawing cross-validation and confidence fusion
[0273] For candidate groups with high conflict, incomplete evidence, or issues involving project consistency, the system generates cross-validation tasks. Validation can involve re-querying the scene graph, reading original attributes, retrieving specification clauses, verifying the device list against the calculation sheet, or comparing versions and cross-page interfaces. Cross-version object matching uses bit number, category, position, attribute, bounding box, and topological neighborhood simultaneously to avoid omissions caused by comparing only file pixels or bit number lists.
[0274] To obtain locatable evidence from specifications, equipment lists, calculation sheets, and manufacturer samples, a document fragment retrieval score is calculated:
[0275] ,
[0276] in, This indicates the retrieval score of the data segment for the query; Indicate candidate data fragments and search questions; Indicates the keyword matching score; Indicates semantic vector similarity; Indicates that the clause number, item, and document type are prior; This indicates the weight of each retrieval item. The retrieval score integrates keyword matching, semantic similarity, item number, item scope, and document type priors, and the returned results also save the file, page number, and fragment range.
[0277] For cross-graph or version-based tasks, calculate the matching cost between the baseline object and the target object:
[0278] ,
[0279] in, This represents the matching cost between the baseline object and the target object; This represents two candidate objects in the previous and next versions; Indicates differences in tag number; Indicates differences in object categories; Indicates differences in location and attributes; Indicates the overlap of the bounding boxes; Indicates topological neighborhood differences; This represents the matching cost weight. The formula integrates tag number, category, position, attribute, bounding box, and neighborhood differences. Unmatched baseline objects become deletion candidates, and unmatched target objects become addition candidates.
[0280] The newly added validation evidence is fused with the original candidate evidence to obtain the confidence level after deducting conflicts and data gaps:
[0281] ,
[0282] in, Indicates the confidence level after candidate group fusion; Indicates the prior confidence level of the candidate group; Indicates the candidate question group; This indicates the independence of the source, the weight of evidence, and the confidence level of each item; Indicates logarithmic probability transformation; This indicates the degree of conflict and data gaps; This represents the penalty coefficient for conflicts and gaps. The formula integrates independent evidence in the log-odds domain and subtracts for conflicts and data gaps. Deterministic tools and independent documents can improve confidence levels, while results from similar sources are discounted.
[0283] Through the above processing, the system can simultaneously verify issues such as data existing but not on the diagram, data appearing on the diagram but not in the data, inconsistent parameters for the same object, cross-page links not pointing back, and object additions, deletions, or modifications in previous versions. The cross-validation results are added to the candidate group as new evidentiary events, while the original candidates are retained, facilitating reviewers' viewing of the conclusion evolution process.
[0284] 4) Feedback and correction, manual confirmation and final conclusion formation
[0285] After cross-validation, the coordinating agent adjusts the candidate states based on the updated support, conflict, completeness, and fusion confidence. Conflicts that can be resolved by the original attributes, geometric retesting, or explicit clauses are automatically corrected; conflicts involving implicit design intent, operational suitability, or data authenticity generate manually confirmed items with evidence from both parties. Manual confirmations, rejections, modifications, and closures are all saved as append events and do not overwrite the original machine records.
[0286] To update candidate confidence based on newly added support, opposition, conflict, and human intervention events, feedback correction is performed:
[0287] ,
[0288] in, Indicates the confidence level of the current round of candidates; This indicates the confidence level after feedback correction; This represents the function that truncates the effective probability range. Indicates support weights and opposition weights; Indicates the degree of conflict and human-caused incidents; This represents the update coefficient for each feedback item. Supporting evidence increases candidate confidence, while opposing evidence and conflicts decrease confidence. Authorized human intervention can confirm or reject the current state. Update results are limited to the effective probability range.
[0289] Candidate states are determined based on the revised confidence level, conflict level, data completeness, and human intervention:
[0290] ,
[0291] in, Indicates the status of candidate questions; Represents the state selection function; This indicates the quantity to be manually confirmed. This indicates the machine verification indication quantity; Indicates the quantity of inference and suggestion; Indicates insufficient data; Indicates the quantity of evidence conflict; This represents the threshold for each state determination. The formula distinguishes between manual confirmation, machine verification, inference suggestions, insufficient data, and conflicting evidence. The state threshold is configured based on the project's risk level; the model cannot directly elevate its own suggestion to a confirmed conclusion.
[0292] After completing the state determination, a final set of questions is generated, preserving historical events:
[0293] ,
[0294] in, This represents the final set of questions following the ruling; This represents a final question; Indicates the problem processing status; This indicates confirmation, verification, suggestion, insufficient information, and conflict. Indicates the history of the problem event; This indicates an event append function; This represents all events related to the current issue. The formula incorporates all issues requiring delivery or further processing into the final set, and appends candidate, validation, and human events to ensure the traceability of the feedback and correction process.
[0295] After the above processing, each final issue has a unique number, target object, evidence level, confirmation status, fusion confidence level, and complete timeline. The report may only display the merged main issue, but the reviewer can expand to view the original conclusions of each agent, supporting or opposing evidence, re-verification results, and manual processing records.
[0296] Through the four steps described above, multi-source candidate issues undergo object-level grouping, evidence quality calculation, cross-document and cross-graph verification, and feedback correction to form a final audit conclusion with a clear state and a complete chain of evidence. This mechanism reduces duplicate issues and contradictory conclusions, and accurately transfers engineering judgments that cannot be automatically resolved to the auditors.
[0297] (6) Audit conclusion positioning, labeling and safe delivery module
[0298] This module is used to convert the final conclusions after adjudication into viewable, locatable, verifiable, and rectifiable engineering deliverables. Based on the evidence of problems, the system generates structured records, drawing location frames, browser view commands, audit reports, and independent annotation files. When a user authorizes modifications to a CAD copy, it employs plan preview, field whitelisting, new file output, and post-write readback verification to prevent the agent from directly overwriting the original deliverables.
[0299] 1) The final review conclusion is generated in a structured manner.
[0300] The system first converts the final set of issues into unified review conclusion records. Each record includes at least the issue claim, target object, issue type, severity, evidence level, confirmation status, fusion confidence level, object evidence, document basis, spatial location, rectification suggestions, and historical events. Cross-map issues can be associated with multiple targets and locations, but are counted as only one logical issue in the project issue statistics.
[0301] To ensure a consistent representation of issues from different sources and scopes, a final review conclusion record is constructed:
[0302] ,
[0303] in, This indicates a final review conclusion; Indicates the conclusion number; Indicate the claim, the goal, and the type of problem; Indicates severity, level of evidence, and status; Indicates the fusion confidence level; Represents the set of evidence and the set of locations; This section indicates rectification suggestions and historical events. The final record saves the identity, claim, target, type, severity, level of evidence, status, confidence level, evidence, location, suggestion, and history. All fields are derived from the ruling.
[0304] Calculate the risk and severity based on the likelihood of the problem occurring, the scope of its impact, and the uncertainty of the evidence:
[0305] ,
[0306] in, This indicates the overall risk of the problem; Indicates the likelihood of a problem occurring; The sub-items are: severity, cost, spread, evidence, and uncertainty. Indicates the weight of each risk component; Indicates the severity level of the problem; This represents the risk level mapping function. The formula incorporates the probability of occurrence, severity of impact, rectification costs, scope of spread, and uncertainty into the risk assessment. High-risk items awaiting confirmation can be prioritized for review but will not be displayed as confirmed errors.
[0307] To prevent issues with no object, no evidence, or no state from entering the report, calculate the delivery conditions for the conclusions:
[0308] ,
[0309] in, Indicate whether the conclusion meets the delivery conditions; This indicates that the field contains an indicator function; Represents the target object; Represents a set of evidence; Represents a set of locations; Indicates confirmation status; This indicates an empty field or an empty set. The formula requires the target object, evidence, location, and state to all exist simultaneously. Records that do not meet the conditions are retained in the internal candidate area; the reporting module must not create a formal question solely based on the natural language summary.
[0310] Through the above processing, the final audit conclusion has unified fields and clear status, which can simultaneously support project statistics, issue screening, manual review, and external system integration. When an issue is manually confirmed or closed, a new status event is added, while the original rule evidence, model evidence, and adjudication records remain unchanged, thereby maintaining the audit responsibility chain.
[0311] 2) Problem location reverse mapping to the original CAD object
[0312] After generating the final conclusion, the system calculates the problem location based on the standard bounding box, insertion point, topological node, and original handle of the evidence object. Evidence that is close to each other and belongs to the same page is merged into a local location box, while evidence that is far apart, spans multiple pages, or spans multiple figures generates multiple location records. Questions that only have model coordinates but no real object are assigned different states to avoid confusion with editable entities.
[0313] To represent a problem that may have multiple drawing locations, we construct grouped location records:
[0314] ,
[0315] in, This represents the set of grouping positions for the problem; Indicates the source drawing file; Indicates the page to which it belongs; Indicates a local positioning box; Represents the original set of handles; Represents a set of discrete location points; This indicates the number of location groups. Each location record contains the source file, drawing page, bounding box, handle set, and discrete points. Multi-page issues do not forcibly merge coordinates from different drawings into a single rectangle.
[0316] For the same location group, a visual bounding box is generated based on the area enclosed by the evidence object:
[0317] ,
[0318] in, A positioning box representing a group of positions; Indicates the function that extends the positioning range; Represents the minimum bounding box function; This represents the set of evidence objects within the location group; Represents the standard bounding box of the evidence object; This represents the margin function determined by the problem type and map size. This indicates the issue type and drawing scale. The formula first takes the minimum outer bound of the evidence object, and then adds the display margin according to the issue type, text height, and drawing scale, so that the positioning frame covers the evidence without excessively obscuring the drawing.
[0319] Reverse map the standard engineering coordinates to CAD coordinates and calculate the coordinate mapping error:
[0320] ,
[0321] in, Represents the CAD coordinates after reverse mapping; Represents standard engineering coordinates; This indicates the inverse transformation of the coordinates of the corresponding drawing; This indicates the coordinate back-and-forth mapping error; Indicates Euclidean distance; This indicates the allowable mapping error threshold. The formula uses the inverse transformation saved in the first module to return the CAD coordinates and verifies the round-trip error. Locations exceeding the threshold are marked as mapping anomalies and cannot be used for automatic annotation or write-back.
[0322] After the above processing, each issue can be located to a specific file, page, world coordinates, bounding box, and original handle. When a user clicks on a report or issue list, the system can directly open the corresponding page and focus on the evidence object; for cross-page issues, the interface displays each location and the relationships between them in sequence.
[0323] 3) Drawing annotation, visual commands, and report content generation
[0324] The system generates commands for object spotting, path drawing, point-of-fact marking, and drawing annotation based on the issue type and evidence level. Different colors, line types, and labels are used for definitive issues, inference suggestions, insufficient data, and conflicting evidence. View commands only affect browser display and do not modify the original drawings. When offline viewing is required, the system can generate rectangles, leader lines, and text annotations in a new review layer and generate DOCX, PDF, Excel, or structured data reports from the same final issue set.
[0325] To make the standard coordinates consistent with the browser view Figure 1 To generate coordinate transformations and view commands:
[0326] ,
[0327] in, This indicates that the browser will display the coordinates; Indicates viewport scaling and translation transformations; This indicates a page cropping transformation; Indicates the inverse transformation of standard coordinates; Represents standard engineering coordinates; The set of visual commands that represent the problem; Represents functions for object spotting, path drawing, and suspicious point marking; This represents the positioning box, path, and set of suspicious points. Standard coordinates are first converted to drawing coordinates, then zoomed and panned through the viewport to obtain the displayed coordinates. View commands include object focus, connecting paths, and discrete suspicious point markers.
[0328] For issues requiring offline CAD viewing, construct annotation objects in a separate review layer:
[0329] ,
[0330] in, This represents a CAD audit annotation object; Indicates an independently reviewed layer; Indicates the geometry of the annotation; Indicates the annotation text; This represents a function that defines the annotation style. Indicates the level of evidence and the status of confirmation; This indicates the source issue number. The formula specifies that the labeled object has an independent layer, labeled geometry, label text, evidence status style, and source issue. The labeled object does not change the entity being audited itself.
[0331] The final set of questions generates reports in various formats, and the completeness of the report content is calculated.
[0332] in, This indicates the audit report; This indicates the report rendering function; Represents the problem grouping function; Represents the final set of problems; Indicates the report template and language region; Indicates the completeness rate of the report content; Indicates the importance of the issue; This indicates the issue delivery conditions. The formula organizes the final issues by project, drawings, severity, and status, and renders the report from a unified record. The report module only allows sorting and wording; it does not allow adding new audit facts.
[0333] After the above processing, the online workbench, offline annotated drawings, and audit reports share the same location, evidence, and confirmation status. Auditors can jump from the report to the drawings and also return to the issue details from the drawing annotations; the report cover also lists the input scope, rules and model versions, data summary, unexecuted tasks, and audit coverage.
[0334] Through the above three steps, the final audit conclusion is transformed into an engineering deliverable with a unified structure, accurate location, visual annotation, audit report, and controlled modification capability.
[0335] Experimental verification
[0336] To verify the effectiveness of this invention in object structure utilization, open semantic auditing, evidence localization, and multi-agent collaboration, existing content constructs an industrial CAD auditing test suite oriented towards embodiments. This suite includes piping schematics, mechanical parts and assembly drawings, architectural drawings, structural drawings, and electrical drawings. The samples include both anonymized actual DWG or DXF files and controlled defect versions created by deleting entities, modifying attributes, breaking line segments, replacing dimensions, changing title blocks, and cross-page targets. ① Process piping diagrams are used to verify tag numbers, pipe connections, valve and instrument relationships, cross-page interfaces, and specification issues; ② Mechanical parts and assembly drawings are used to verify dimensions, tolerances, roughness, materials, hole positions, and version differences; ③ Architectural drawings, structural drawings, and electrical drawings are used to verify frame text, room or component relationships, detail references, loop equipment numbering, data verification, and cross-drawing relationships. Controlled modifications record the target handle, values before and after modification, and injection type, and the file is re-parsed after modification to confirm consistency between the drawing and the object; manual annotations include at least the problem type, severity, drawing page, object handle or bounding box, applicable rules, and rectification description, and retain the pending confirmation status for connection intentions or specification issues that can only be confirmed through engineering judgment.
[0337] The implementation plan collected 480 drawings from 72 projects or independent drawing sets. 300 drawings were used for project-specific configurations, threshold development, and model debugging, while 180 were used as an isolated test set. To avoid leakage of similar drawings from the same project, drawings from the same project were only included in one data partition. Test drawings were annotated by three personnel with CAD experience and manually reviewed and approved, ultimately resulting in 1420 real-world issues and 3580 normal review units, totaling 5000 decision units. Real-world issues covered eight categories: numbering and title blocks, annotations and text, equipment and documentation, geometric dimensions, pipeline connections, standard application, cross-drawing interfaces, and version changes. Cross-drawing issues could be associated with more than two drawings, but were counted as only one issue.
[0338] To compare different technical approaches, five methods were selected and tested under uniform conditions: ① "Image + OCR" renders CAD drawings as 300 dpi images, outputting problems through symbol detection, character recognition, and fixed post-processing; ② "Fixed Rules" directly parses CAD objects and executes tag number, layer, title block, dimension, and connection rules, but does not use a language model or rich neighborhood; ③ "Raster Single Agent" provides drawing images, OCR text, and review targets to a single language model, without providing original handles, scene diagrams, or deterministic tools; ④ "Scene Diagram Single Agent" uses the object parsing and layered scene diagrams of this invention, but all tasks are completed serially by a single agent, without employing multi-agent division of labor, dual-track adjudication, or checklist-based routing capabilities; ⑤ "Invention Method" uses all technical modules.
[0339] All methods use the same test set, problem type range, project data visibility, and hardware environment. The three methods requiring a model use the same DeepSeek model version, output length, and temperature value, and are run with five different random seeds. Repeated runs following fixed rules maintain consistent results. Object parsing and the initial DWG conversion are included in the single-image review time; cache hits are not counted repeatedly. Pixel bounding boxes that cannot be mapped to objects through standard coordinate transformation are considered positioning errors. This invention first executes a deterministic caliper, then generates tasks according to map type, checklist, and project scope. A single rich neighborhood can contain a maximum of 320 main objects, with a maximum concurrency of 4 and a maximum nesting depth of 3, and forces the model inference to remain in a pending confirmation state. Test maps are categorized into low, medium, and high complexity based on the number of primitives, object density, layer and block references, cross-map edges, and problem type.
[0340] The evaluation was conducted using the following metrics: Problem Detection Rate (DR), Effective Problem Rate (VR), Location Hit Rate (LH), Evidence Completeness Rate (EC), Review Coverage Rate (ACov), and Average Review Time. Specifically, DR represents the proportion of genuine problems detected by the system; VR represents the proportion of system-outputted problems that were confirmed as valid after manual review; LH represents the proportion of problem conclusions that accurately match the handles or location boxes of labeled objects; EC represents the proportion of conclusions that simultaneously include object source, rule or clause basis, location information, evidence description, and confirmation status; ACov represents the proportion of preset check items, drawing objects, and cross-drawing relationships that the system effectively reviewed; and Average Review Time includes the processes of drawing parsing, scene graph construction, agent reasoning, dual-track evidence adjudication, and report generation.
[0341] Table 1. Comparison of data from different methods under six major project review indicators.
[0342]
[0343] The experimental results are shown in Table 1. Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 As shown, the problem detection rate, effective problem rate, location hit rate, evidence completeness rate, and review coverage rate were 91.8%, 93.6%, 98.2%, 96.4%, and 94.7%, respectively, which are generally better than the other four methods. The average time of 41.6 seconds consists of 8.4 seconds for DWG conversion and object parsing, 6.7 seconds for scene graph and indexing, 4.3 seconds for deterministic caliper, 18.9 seconds for multi-agent reasoning, and 3.3 seconds for evidence adjudication and report organization.
[0344] Experimental results are as follows Figure 8 As shown, in the complexity test, the problem detection rates of this invention on low, medium, and high complexity drawings were 95.5%, 92.8%, and 89.7%, respectively. The detection rate decreased by 5.8 percentage points for high complexity compared to low complexity, while it decreased by 14.2 and 18.1 percentage points for grid-based single-agent and fixed-rule drawings, respectively. The relatively low performance at high complexity is mainly due to the lack of operating conditions or manufacturer parameters for some items; the system marks these as insufficient data rather than forcibly determining violations, which aligns with the engineering credibility boundaries set by the existing content.
[0345] Overall experimental results show that the image + OCR method can recognize clear text and some standard symbols, but it is difficult to preserve the original layers, object handles, block attributes, and accurate topology; the fixed rule method is suitable for handling deterministic matters where attributes can be directly read or residuals can be calculated, but it is difficult to cover open connections, canonical interpretations, and cross-data issues; the grid single agent improves the detection capability of complex semantic problems, but still suffers from errors in locating devices with the same name, misidentification of legends, and improper evidence status; the scene graph single agent demonstrates the role of hierarchical object representation and rich neighborhood in improving the location hit rate, evidence completeness rate, and review coverage rate, but when a single context undertakes all tasks, there are still omissions, duplicate conclusions, and conflicting conclusions in cross-graph and data tasks.
[0346] This invention further improves the overall review quality through multi-agent division of labor, dual-track evidence of determinism and inference, conflict adjudication, and capability routing, and avoids duplicate review during the report aggregation stage by parallelizing tasks. Combining existing embodiments, statistical tests, and controlled write-back results, this invention can improve the detection capabilities of complex connections, specification data, and cross-version issues while maintaining object location and false alarm control capabilities, and ensures that model inference does not exceed engineering confirmation boundaries, thereby improving the automation, accuracy, traceability, and engineering application security of industrial CAD drawing review.
[0347] Example 2
[0348] This embodiment provides a multi-agent review system for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0349] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device, the aforementioned method for multi-agent review of industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0350] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement various instructions; the computer-readable storage medium being configured to store multiple instructions adapted for loading and execution by the processor of the aforementioned multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration.
[0351] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration, characterized in that, include: Industrial CAD drawing access and traceable object parsing, including drawing file reception and parsing scope determination; Unified modeling of native primitive objects; Geometric coordinates, drawing frames, and scale standardization; Standard object library and traceability mapping output; Based on the parsing results, addressable primitives and scene graphs are constructed, including primitive node feature and spatial index construction; primitive relationships, annotation orientation and topological edge generation; engineering semantic object recognition and hierarchical aggregation; Addressable primitives and scene graph output; Review rule context generation and multi-agent task scheduling, including review rule and checklist structuring; Rule fit calculation and review context filtering; review task set generation and dependency relationship construction; task agent matching and controlled scheduling; Professional audit agent execution and dual-track evidence generation, including professional audit task context construction; professional agent audit reasoning and tool verification; Deterministic inference of dual-track candidate problem expression; summarization of multi-task review results and scope verification; Multi-source result collaborative adjudication and cross-graph feedback correction, including candidate problem association grouping and duplicate resolution; Calculation of evidence support, conflict level, and data completeness; cross-document and cross-drawing cross-validation and confidence fusion; Feedback and correction, manual verification, and final conclusion formation; Audit conclusion location, annotation, and secure delivery, including the structured generation of the final audit conclusion; reverse mapping of issue locations to original CAD objects; Drawing annotations, visual commands, and report content generation.
2. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 1, characterized in that, The industrial CAD drawing access and traceable object parsing specifically includes: first, constructing a drawing input set and calculating the format validity and task input completeness rate of each file; for files that pass basic verification, performing controlled transformation and object-level parsing, and outputting drawing records with file summary, format status, transformation source, and parsing log; after determining the drawing input range, using a parser to traverse the model space, layout space, block definition, block reference, and attribute reference, reading line segments, circles, arcs, polylines, splines, fills, text, dimensions, leader lines, tables, and proxy objects, and, based on a unified object record, splitting geometric information into continuous control points and type-specific parameters; finally, for use in relation calculation and semantic recognition, further constructing fixed-dimensional object input features by classifying object type, geometry, and layers. Style, attribute, and region location encoding are used as unified input features to form a consistent and expandable object record for each CAD entity. To address scale deviations and page number mismatches that occur when directly using raw coordinates to perform distance, connection, and cross-drawing comparisons, CAD header variables, viewport scale, drawing frame boundaries, and title block information are read to establish a unified engineering coordinate transformation for each drawing and recalculate the object's spatial extent. This ensures that the original object obtains coordinates, length, and bounding box at a unified scale and has a clear or unconfirmed drawing page affiliation. After completing object resolution and geometric standardization, an internal identifier is generated for each object that does not conflict across drawings and is stable when repeatedly resolved within the same file. A two-way mapping is established between the internal identifier, file summary, original handle, drawing page, and standard bounding box, while recording the scale and resolution quality of unknown objects.
3. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 2, characterized in that, The construction of addressable primitives and scene graphs based on the parsing results specifically includes: first, converting each standard object into a primitive node, with each node carrying object type, standard geometry, layer, style, text attributes, drawing page, and source reliability; to support neighborhood queries in large-scale drawings, constructing a spatial index based on the standard bounding box and insertion point, and registering the drawing frame, title block, legend, equipment list, and main drawing area as region objects, so that text with the same name receives different semantic weights depending on its region; then, calculating intersection, adjacency, parallelism, perpendicularity, collinearity, concentricity, containment, and endpoint connection relationships based on standard geometry; for leaders and multiple leaders, extracting the arrow ends and text ends, and determining target candidates by combining distance, object category, layer, and text mode; for pipeline objects, performing tolerance clustering on the endpoints and generating topological edges with original handles; After obtaining primitive nodes and relation edges, candidate groups are generated by combining block definitions, attribute keys, tag patterns, adjacent text, object relationships, and project conventions. Then, the contribution of each primitive to the engineering object is calculated through attention weights. Finally, primitive nodes, engineering semantic nodes, spatial relationships, annotation relationships, topological relationships, constraint relationships, composition relationships, and cross-graph knowledge relationships are uniformly organized into an addressable hierarchical scene graph. Lower-level nodes provide precise geometry and raw handles, while higher-level nodes express device roles, system affiliation, parent-child relationships, specification clauses, and version relationships. Each layer is connected through stable mapping, enabling audit rules and agents to query according to the required evidence level.
4. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 3, characterized in that, The aforementioned review rule context generation and multi-agent task scheduling specifically include receiving enterprise checklists, mapping specifications, project agreements, and user natural language objectives. Each review item is broken down into review dimensions, applicable map types, target objects, necessary inputs, judgment modes, available tools, result status, and manual conditions. Items that can be directly judged through fields, quantities, dimensions, or topology are marked as deterministic capabilities; those related to connection intent, specification applicability, and experiential risks are marked as intelligent preliminary judgment or manual confirmation capabilities. To enable the natural language review items to be executed by the system, a structured rule record is first constructed. Based on the rule record, the weighted completeness of the inputs required for the review item is calculated. Finally, the processing mode for the review item is selected based on the review content, available capabilities, execution risks, and costs. After obtaining the structured inspection items, the rule fit is calculated by comprehensively considering map type matching, target object hit, completeness of necessary input, tool usability, and historical reliability. For rules that pass the screening, rich neighborhoods are constructed around the target object, tag number, user selection area, or cross-page interface, and the context range is adjusted according to the map sheet scale, path coverage, and object density. After the rule fit and context screening are completed, the audit task is generated according to the scope of drawing level, drawing page level, object level, system level, cross-drawing level, and cross-document level. Each task records the target object, professional field, inspection basis, input reference, risk level, expected output and permitted tools, and calculates priority based on risk, hard error attributes, user attention and expected cost; after forming a set of audit tasks, the matching degree is calculated based on the domain capabilities, callable tools, processing scope, permission range, current load and historical reliability of each professional intelligent agent.
5. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 4, characterized in that, The professional auditing agent executes and generates dual-track evidence, including extracting primitives, semantic objects, spatial neighborhoods, topological paths, annotation relationships, inspection rules, and project data fragments from the hierarchical scene graph according to the task objectives. The context is stably sorted according to page, coordinates, object type, and topological level, and the original facts, rule derivation, semantic inference, document references, and truncated ranges are clearly identified to help the agent distinguish between known evidence and unprovided information. Based on the task type, it performs format verification, connectivity analysis, specification retrieval, device data verification, cross-graph comparison, or version impact analysis. The language model is responsible for understanding open design intentions and organizing tool calls, while the deterministic tool is responsible for returning the number of objects, attribute values, geometric distances, topological paths, and the original text of the clauses. To differentiate between different proof strengths, a deterministic evidence track and an inferential evidence track are established. Number, quantity, version, object attributes, dimensional residuals, and provable topological results enter the deterministic track when the input is complete and the decision boundaries are clear; connection intent, normative interpretation, empirical risks, and visual doubts enter the inferential track, and manual confirmation marks are mandatory. After multiple professional intelligent agents complete their tasks, candidate issues are summarized according to drawings, pages, inspection items, systems, and cross-drawing scopes, while retaining task status, reasons for non-execution, and operation logs.
6. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 5, characterized in that, The multi-source result collaborative adjudication and cross-graph feedback correction include describing the same problem from the perspectives of rules, connections, norms or cross graphs for different agents, giving claims that are similar in text but different in reality on adjacent objects, comprehensively calculating the candidate association degree by integrating the original handle set, standard bounding box, problem type, rule clause, check item and claim semantics, and then forming problem groups by the connected components of the association graph; After candidate grouping is completed, support weights are calculated based on evidence level, source reliability, source independence, object hit rate, and citation completeness. Conflict degree is calculated for mutually exclusive categories, opposing conclusions, numerical out-of-limits, and normal relationships. Input gaps are calculated for groups lacking drawings, clauses, or external parameters. For candidate groups with high conflict, incomplete evidence, or involving project consistency, cross-validation tasks are generated. Validation is used to re-query scene diagrams, read original attributes, retrieve specification clauses, verify equipment lists and calculation sheets, or compare versions and cross-page interfaces. After cross-validation, the coordinating agent adjusts the candidate status based on the updated support, conflict degree, completeness, and fusion confidence. Conflicts that can be resolved by original attributes, geometric retesting, or explicit clauses are automatically corrected. Conflicts involving implicit design intent, operational suitability, or data authenticity generate manual confirmation items with evidence from both sides. Manual confirmation, rejection, modification, and closure are all saved as append events and do not overwrite the original machine records.
7. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 6, characterized in that, The process of locating, labeling, and securely delivering audit conclusions includes first converting the final set of issues into unified audit conclusion records. Each record must include at least the issue claim, target object, issue type, severity, evidence level, confirmation status, fusion confidence level, object evidence, document basis, spatial location, rectification suggestions, and historical events. After generating the final conclusion, the issue location is calculated based on the standard bounding box, insertion point, topology node, and original handle of the evidence object. Evidence that is close to each other and belongs to the same drawing page is merged into a local location box, while evidence that is far apart, spans multiple pages, or spans multiple drawings generates multiple location records. Questions with only model coordinates and no actual object are assigned different statuses. Finally, based on the issue type and evidence level, commands for object spotting, path drawing, question mark marking, and drawing annotation are generated. Deterministic issues, inferred suggestions, insufficient data, and evidence conflicts are assigned different colors, line types, and labels. View commands only affect browser display and do not modify the original drawings. When offline viewing is required, rectangles, leader lines, and text annotations are generated in a new audit layer, and DOCX, PDF, Excel, or structured data reports are generated from the same final issue set.
8. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 7, characterized in that, The problem location is mapped inversely to the original CAD object, including constructing grouped location records to represent the possibility that a problem may have multiple drawing locations: in, This represents the set of grouping positions for the problem; Indicates the source drawing file; Indicates the page to which it belongs; Indicates a local positioning box; Represents the original set of handles; Represents a set of discrete location points; Indicates the number of location groups; each location record contains the source file, page, bounding box, handle set, and discrete points; for the same location group, a visual positioning box is generated based on the bounding area of the evidence object: in, A positioning box representing a group of positions; Indicates the function that extends the positioning range; Represents the minimum bounding box function; This represents the set of evidence objects within the location group; Represents the standard bounding box of the evidence object; This represents the margin function determined by the problem type and map size. Indicate the problem type and drawing scale; reverse map standard engineering coordinates to CAD coordinates and calculate the coordinate mapping error: in, Represents the CAD coordinates after reverse mapping; Represents standard engineering coordinates; This indicates the inverse transformation of the coordinates of the corresponding drawing; This indicates the coordinate back-and-forth mapping error; Indicates Euclidean distance; This indicates the allowed feedback error threshold.
9. The multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 8, characterized in that, The generation of drawing annotations, visual commands, and report content includes generating coordinate transformations and view commands to ensure consistency between standard coordinates and browser views. in, This indicates that the browser will display the coordinates; Indicates viewport scaling and translation transformations; This indicates a page cropping transformation; Indicates the inverse transformation of standard coordinates; Represents standard engineering coordinates; The set of visual commands that represent the problem; Represents functions for object spotting, path drawing, and suspicious point marking; This represents the location box, path, and set of suspicious points; for issues requiring offline CAD viewing, it constructs annotation objects in a separate review layer: in, This represents a CAD audit annotation object; Indicates an independently reviewed layer; Indicates the geometry of the annotation; Indicates the annotation text; This represents a function that defines the annotation style. Indicates the level of evidence and the status of confirmation; Indicates the source issue number; generates reports in various formats from the final issue set, and calculates the report content completeness rate: , in, This indicates the audit report; This indicates the report rendering function; Represents the problem grouping function; Represents the final set of problems; Indicates the report template and language region; Indicates the completeness rate of the report content; Indicates the importance of the issue; Indicates the delivery conditions for the problem.
10. A multi-agent review system for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration, executing the multi-agent review method for industrial CAD drawings based on addressable hierarchical scene graphs and dual-track evidence collaboration as described in claim 1, characterized in that... include: The data acquisition module is configured to handle industrial CAD drawing access and traceable object parsing, including receiving drawing files and determining the parsing scope; unified modeling of native primitive objects; and standardization of geometric coordinates, drawing frames, and scales. Standard object library and traceability mapping output; The primitive construction module is configured to construct addressable primitives and scene graphs based on the parsing results, including primitive node feature and spatial index construction; primitive relationship, annotation pointing and topological edge generation; engineering semantic object recognition and hierarchical aggregation; and output of addressable primitives and scene graphs. The task scheduling module is configured to generate audit rule contexts and schedule multi-agent tasks, including the structuring of audit rules and checklists. Rule fit calculation and review context filtering; review task set generation and dependency relationship construction; task agent matching and controlled scheduling; The evidence generation module is configured to perform professional review agent execution and dual-track evidence generation, including professional review task context construction; professional agent review reasoning and tool verification; deterministic inference of dual-track candidate question expression; and multi-task review result aggregation and scope verification. The feedback correction module is configured to perform multi-source result collaborative adjudication and cross-graph feedback correction, including candidate issue association grouping and duplication resolution; calculation of evidence support, conflict degree and data completeness; cross-document and cross-graph cross-validation and confidence fusion; Feedback and correction, manual verification, and final conclusion formation; The delivery module is configured to locate, label, and securely deliver audit conclusions, including the structured generation of final audit conclusions and the reverse mapping of problem locations to original CAD objects. Drawing annotations, visual commands, and report content generation.