Boiler design file intelligent identification system and method based on multi-modal knowledge graph

The intelligent identification system built through multimodal knowledge graph solves the problems of low efficiency and strong subjectivity of manual identification of boiler design documents, realizes fast and accurate automated identification, reduces the missed review rate, and improves the safety and reliability of boiler design.

CN120496115APending Publication Date: 2025-08-15SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510532937.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The identification of existing boiler design documents mainly relies on manual labor, and there are problems of low efficiency and strong subjectivity, which leads to inconsistent identification results and easy to miss review.

Method used

Using intelligent identification methods based on multimodal knowledge graphs, automatic identification of boiler design files is achieved through data acquisition, multimodal feature extraction and knowledge graph construction. This method includes image feature extraction, text feature extraction and multimodal fusion, to build a knowledge graph with engineering semantics and reasonable capabilities, and to determine whether the design points meet the specifications.

Benefits of technology

It greatly improves the appraisal efficiency, reduces the missed review rate, ensures the accuracy and reliability of the appraisal results, reduces labor costs, promotes the intelligent development of boiler design, and improves safety and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496115A_ABST
    Figure CN120496115A_ABST
Patent Text Reader

Abstract

The invention discloses a boiler design file intelligent identification system and method based on a multi-modal knowledge graph. The method comprises the steps that S1, a data collection module is used for collecting a large number of existing boiler design files as a data set; s2, extracting information on the boiler design drawing in the data set by using a multi-modal feature extraction module to form a text-image multi-modal joint feature vector; s3, constructing a multi-modal boiler design knowledge graph with engineering semantics, graph structure association and reasoning capability; and S4, inputting the extracted key point information on the boiler design drawing to be examined into the knowledge graph model, and judging whether each design point meets the design specification requirements of the boiler drawing or not. According to the method, the identification period is greatly shortened, the working efficiency is improved, and the missed examination rate is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of boiler design document identification, and specifically relates to a boiler design document intelligent identification method and system based on a multimodal knowledge graph. Background Art

[0002] Boiler design documents must be submitted to a qualified inspection agency for evaluation before production can begin. Existing technology primarily relies on manual evaluation of boiler design documents. This process is subject to significant subjectivity and limitations, often limited by factors such as the appraiser's expertise, work experience, and fatigue. This process suffers from the following drawbacks: 1) Low efficiency: The average evaluation time for a single drawing is 2-5 hours; 2) High subjectivity: Different appraisers have varying understandings of the standards, leading to varying degrees of disagreement in the evaluation results. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for intelligent identification of boiler design documents based on multimodal knowledge graph to solve the above technical problems.

[0004] In order to solve the above technical problems, the present invention adopts the following solutions:

[0005] The intelligent identification method of boiler design documents based on multimodal knowledge graph includes the following steps:

[0006] Step S1: using a data acquisition module to collect a large number of existing boiler design files as a data set, the boiler design files including but not limited to boiler design drawings, technical document texts, and three-dimensional model data;

[0007] Step S2: using a multimodal feature extraction module to extract information on the boiler design drawings in the dataset to form a text-image multimodal joint feature vector;

[0008] Step S3: Using the text-image multimodal joint feature vector output in step S2 as input, it is structured and mapped into entity nodes in the heterogeneous graph. Semantic connections are established with the structural constraint information extracted from the boiler design standard specification, thereby constructing a multimodal design knowledge graph with engineering semantics, graph structure association, and reasoning capabilities.

[0009] Step S4: Input the extracted key point information on the boiler design drawings to be reviewed into the knowledge graph model to determine whether each design point meets the requirements of the boiler drawing design specifications; point out the problems in the boiler design and explain the reasons, and generate a review report.

[0010] Further optimization, in step S1, the boiler design drawing image is obtained by a high-resolution scanner, and the technical document text is obtained from the electronic document or paper document after OCR recognition; the three-dimensional model data is generated by professional three-dimensional modeling software or collected by three-dimensional laser scanning equipment.

[0011] Further optimization, in step S2, the multimodal feature extraction module includes an image feature extraction unit and a text feature extraction unit, and the multimodal feature extraction method specifically includes:

[0012] Step S2.1: Image Feature Extraction: First, the HED model and Canny edge detection algorithm are used to extract overall edge information from the design drawing, preliminarily determining the drawing's scope and boiler structure. Next, a U-Net model is used to identify the boiler's contour. By adding an attention mechanism, the model's focus on key boiler components (the boiler body, pipes, valves, etc.) is enhanced, improving the accuracy of boiler contour extraction. The attention mechanism dynamically adjusts the model's attention weights for different regions, ensuring that key component features are not overlooked. Each extracted graphic component is then encoded as a high-dimensional image feature vector, retaining its geometric position information and structural feature parameters, such as circles, ellipses, connecting lines, and dimension lines, to facilitate positional alignment and semantic matching with text features.

[0013] In the application, the HED model was primarily used to extract edge information for complex components, such as the irregular contours of the boiler body, curved sections of pipes, and the complex structure of valves. The Canny algorithm was used to supplement the extraction of simple, regular edges, such as those of straight pipes and circular flanges. This combined approach ensured that all critical edge information in the drawing was extracted, providing high-quality data support for subsequent contour recognition and dimensional measurement.

[0014] Step S2.2: Text feature extraction: Based on OCR technology, text areas are detected from the boiler design drawings and converted into editable text data. The text information in the drawings is accurately extracted, including key content such as boiler component names, material types, and dimension annotations. Then, using the improved BERT model, the text information extracted by OCR is annotated and encoded into high-dimensional semantic features to form a text feature vector. At the same time, the spatial position information of the text in the drawing is retained as the text input for multimodal fusion.

[0015] Specifically, the optimized OCR technology is used to perform image processing and pattern recognition on boiler design drawings, detect text areas and convert them into editable text data, accurately extract key information such as component names, material types, dimension markings, and effectively handle complex backgrounds, tilted text, and blurred characters, thereby improving the accuracy and robustness of text recognition.

[0016] The advantages of the BERT model lie in its good semantic expression capabilities and powerful transfer learning characteristics. It is suitable for scenarios where professional field corpus is relatively scarce, such as boiler design drawings. It can quickly adapt to the terminology distribution and expression style in industrial drawings through the driving force of terminology vocabulary and industry background knowledge.

[0017] Step S2.3: Multimodal fusion: First, the image feature vector and text feature vector are mapped to the same semantic space and aligned. Taking the boiler body as an example, the contour features are aligned with the corresponding text features such as component names and dimension annotations to ensure consistency of multimodal information. Then, the aligned image feature vectors and text feature vectors are sequentially spliced to form a multimodal feature sequence. For example, the contour feature vector of the boiler body is sequentially spliced with the corresponding text feature vectors of component names and dimension annotations to form a multimodal feature sequence. The self-attention mechanism of the Transformer multimodal fusion module is then used to calculate the correlation weights between each feature vector and dynamically adjust the feature representation. For example, the contour features of the boiler body will interact with the text features of component names and dimension annotations, and the correlation weights between them will be calculated, thereby enhancing the semantic information related to the boiler body. The multi-layer Transformer encoder performs nonlinear transformations and interactions on multimodal feature sequences, extracting higher-level semantic representations layer by layer, understanding the structural relationships between components, and fusing multimodal features to generate a unified semantic representation that incorporates both image and text information. This accurately describes the semantic content of boiler design drawings, thereby achieving a comprehensive understanding of the boiler design drawings. This fusion approach not only compensates for the shortcomings of single-modal information but also, through the complementarity between features, further improves the comprehensiveness and accuracy of feature extraction.

[0018] Further optimization, in step S2.3, the image feature vector and the text feature vector are mapped to the same semantic space, aligned and spliced, specifically including:

[0019] Step S2.3.1: Perform normalized spatial embedding based on the spatial position information of each text annotation and graphic component in the drawing. Specifically, the system uniformly maps all drawing coordinates to the interval [0, 1], generating a 128-dimensional position encoding vector. This vector not only contains the center point coordinates but also encodes spatial structural information such as the shape, area ratio, and directionality of the graphic element, thus providing a high-precision position information foundation for image and text alignment.

[0020] Step S2.3.2: Use the nearest neighbor search method to screen all image-text pairs based on the distance or overlap between the spatial position codes, and screen out image-text combinations that are physically adjacent on the drawing and logically semantically associated. Specifically, the K-nearest neighbor algorithm is used to determine a set of candidate image-text matching pairs. This method can quickly screen out image-text combinations that are physically adjacent on the drawing and logically may have semantic associations, laying the foundation for subsequent semantic fusion. For such candidate image-text pairs, prompt engineering technology is used to construct refined prompt templates to guide large language models to analyze whether there is a semantic association between images and texts. Prompt engineering technology allows large models to exert their cross-semantic understanding and reasoning capabilities through the method of "setting context + preset intention" to automatically identify potential connections between images and texts.

[0021] Step S2.3.3: Retain the image-text matching pairs with high semantic correlation to construct the fused text-image joint feature vector, which is specifically expressed as follows:

[0022]

[0023] Where: V text Represents the text semantic vector; V image Represents the graph structure feature vector; V position represents the position encoding vector; Represents vector concatenation.

[0024] In addition, in order to capture the potential cross-modal implicit connections between images and text, the attention mechanism is introduced in the multimodal feature fusion stage. The specific implementation method is: when constructing the joint feature vector V joint Previously, the system introduced a weighted attention layer for the image-text candidate pairs to model the correlation between different tokens in the text semantic vector, such as "material", "thickness", "Q345R", etc., and the local area features of the image.

[0025] For example, the weighted fusion expression is as follows:

[0026]

[0027] Among them, α i Attention weights represent the degree of attention each text token pays to the target graphic area. The system is trained to enhance its perception of specialized terminology (for example, assigning higher weights to "pressure element," "shell," and "weld"). This mechanism enables the system to highlight key semantic points while deemphasizing irrelevant information when fusing features, thereby improving the discriminability of multimodal vectors.

[0028] Further optimization, the step S3 of constructing a dynamic knowledge graph specifically includes:

[0029] Taking the text-image multimodal joint feature vector output in step S2 as input, it is structurally mapped into entity nodes in the heterogeneous graph and semantically connected with the structural constraint information extracted from the standard specification clauses, thereby constructing a multimodal design knowledge graph with engineering semantics, graph structure association and reasoning capabilities;

[0030] Step S3.1: Basic vector fusion and preparation: Summarize the BERT encoded semantic feature vector, graph structure vector, and normalized position encoding vector generated in step S2, and use the attention weighting mechanism to fuse these three types of vectors to obtain a unified embedding representation, providing basic semantic input for graph node modeling.

[0031] Step S3.2: Construct entity node:

[0032] Step S3.2.1: Semantic type determination: A multi-classification model based on XGBoost, with its excellent nonlinear expression and feature selection capabilities, processes each nested structure that integrates text, graphics, and position semantics in the boiler design drawings, performs semantic type determination, and outputs predefined node types such as "boiler components", "structural parameters", "material types", and "drawing annotations".

[0033] Step S3.2.2: Attribute parsing and encapsulation: Call the corresponding attribute parsing template based on the node category, perform structured encapsulation on the semantic information, spatial parameters and original extracted metadata in the embedded vector, and generate standardized fields such as component type, thickness, unit, location and drawing number.

[0034] Step S3.2.3: Node storage and annotation: Format the structured nodes according to the JSON-LD semantic modeling specification, write them to the database by calling NebulaGraph's HTTP interface, and complete type annotation and index registration.

[0035] Step S3.3: Generate semantic edge relationships: The generation of semantic edge relationships is accomplished through the collaborative efforts of rule-driven and semantic similarity-driven mechanisms. The rule-driven path is based on predefined relationships within engineering templates, such as establishing "used materials" edges between material annotations and components, and "design parameter" edges between dimension annotations and components. The semantic similarity-driven path uses Sentence-BERT to encode node content, and uses vector matching to select entity pairs with cosine similarity exceeding a threshold as potential edge candidates. It then combines spatial proximity, drawing annotation distance, and other conditions to determine whether to establish semantic edges such as "structural dependency," "functional attribution," and "drawing annotation pointing."

[0036] Step S3.4: Standard clause processing and integration

[0037] Step S3.4.1: Acquisition and preprocessing of standard clauses: Using the Boiler Safety Technical Regulations TSG11-2020 and its supporting specification documents as data sources, LlamaIndex is used to perform PDF segmentation preprocessing on the standard documents, and the AIOpenAI text-embedding-ada-002 model is used to build a clause-level vector index library;

[0038] Step S3.4.2: Clause Extraction and Structuring: For the nodes generated in the drawing, construct an embedding query and retrieve the semantically closest clause paragraphs through the LlamaIndex retrieval interface. Input the retrieved paragraphs into the DeepSeek-LLM7B-InstructA1 model deployed in the local Ollama container environment. Using prompt engineering, construct standardized prompts to extract structured triples.

[0039] Step S3.4.3: Fusion of Standard and Design Nodes: Perform regularization verification, logical consistency check, and unit normalization on the extracted results. Use Cypher query language to inject the newly generated standard nodes and their constraint relationship edge structures into the NebulaGraph graph to achieve semantic connection between the design nodes and the standard specifications.

[0040] Step S3.5: Compliance reasoning model construction and training

[0041] Step S3.5.1: Model selection and implementation: We select the Relational Graph Attention Network (R-GAT) as the main inference model and build the model structure based on the PyTorchGeometric framework. This model supports edge type awareness, multi-head attention mechanism, and node embedding updates, and is used to learn the weight distribution of multi-relation paths in heterogeneous graphs.

[0042] Step S3.5.2: Training sample construction: Select the node of the component to be reviewed from the archived boiler drawing dataset as the center, expand the two-hop adjacent nodes (parameter, material, standard clause nodes, etc.) outward to form a local heterogeneous subgraph, and use the expert review conclusion (compliant / non-compliant) as the classification label to construct supervised training samples;

[0043] Step S3.5.3: Model training and optimization: During the training process, the cross-entropy loss function is used, and the graph neural network learns the importance of semantic paths through multiple rounds of message passing, and optimizes the node representation under the edge type weighting mechanism.

[0044] Further optimization adopts graph structure perturbation and data enhancement strategies, including node attribute perturbation, graph structure pruning, historical version difference annotation and multi-region sampling mechanism, to improve the reasoning stability and generalization ability of the inference model under different boiler types, structural layouts and design styles.

[0045] Further optimization, step S3 also includes boiler design specification standard update and graph reconstruction, specifically: First, use the standard automatic acquisition and structured update module based on the Scrapy framework to crawl the standard updates, clause revisions and technical notifications from platforms such as the State Administration for Market Regulation. Use the OCR engine to identify standard documents, combine with HanLP to complete syntax analysis, field extraction and keyword matching, and extract clause numbers, constraint logic and applicable condition structure fields; then, write the new clauses into the NebulaGraph graph after structured processing, establish "substitution relationship" edges pointing to the old standard clauses, trigger the dependency path update and embedding reconstruction of related design nodes, realize the automatic coordination of graph version evolution and model adaptation, and ensure that the system is synchronized with the latest industry standards.

[0046] Further optimization, in step S4, the intelligent comparison unit uses a rule engine to input the constructed sub-graph of the drawing to be reviewed into the trained R-GAT model, output the node classification results and their corresponding attention path weights, and evaluate whether the key node structure in the drawing to be reviewed meets the constraints of standard specifications; and through the graph visualization module, dynamically display the key path and its semantic logic used by the R-GAT model in judgment and prediction, providing path-level diagnostic basis for manual review.

[0047] Further optimization, in step S4, different judgment conditions are corresponding to different key points extracted from the boiler design drawings, and the problem points in the boiler design drawings are divided into problem points that are vetoed and problem points that allow a certain error;

[0048] Among them, the issues and criteria for veto include:

[0049] 1) The safety water level marking is incorrect or missing

[0050] Judgment conditions: The lowest safe water level of the boiler is not clearly marked in the design drawing or the lowest water level is lower than the highest fire limit; the lowest water level of the shell boiler should be 100mm higher than the highest fire limit (75mm for inner diameter ≤1500mm).

[0051] 2) Lack of emergency water discharge device

[0052] Judgment conditions: The power station boiler drum is not equipped with an emergency water discharge device, or the discharge pipe is located below the minimum safe water level;

[0053] 3) The door hole design does not meet the maintenance requirements

[0054] Judgment conditions: The number of manholes, handholes, and headholes does not meet the minimum requirements or the arrangement is not convenient for cleaning and maintenance. If the inner diameter of the boiler drum is ≥800mm and no manhole is set, it is a direct rejection item;

[0055] 4) The key T-joint adopts overlapping structure

[0056] Judgment conditions: The T-joint at the flue gas scouring area does not adopt a full-penetration butt joint structure with groove processing, and there is an overlap connection;

[0057] 5) The weld does not meet the minimum distance requirements and there is high stress concentration

[0058] Judgment conditions: The center line spacing between the drum and furnace welds is less than 3 times the thickness of the steel plate and less than 100mm, which is prone to structural defects and is directly judged as unqualified.

[0059] 6) Welding quality and non-destructive testing

[0060] The pass rate of non-destructive testing (X-ray, ultrasonic) of welds must be 100%, and the film spot checks must cover key areas (weld intersections, T-joints).

[0061] Problem points and judgment conditions that allow a certain degree of error include:

[0062] 1) Detail error of water level gauge scale line:

[0063] Allowable error: ±2mm does not affect safety judgment, and the system prompts manual review.

[0064] 2) The manhole and handhole are slightly smaller in size but have a reasonable structure:

[0065] Allowable error: Due to structural limitations, if the height or diameter is smaller than the standard but has the possibility of entry, it is considered a tolerance design and requires manual judgment;

[0066] 3) Pipe hole layout margin error

[0067] Allowable error: The edge distance or hole spacing is slightly smaller than the standard but not overlapping. When the hole diameter is <60mm, the edge distance requirement can be reduced by 5%.

[0068] 4) Expansion indicator marking deviation

[0069] Permissible error: For Class A boilers, an offset of the expansion indicator of less than 10% is tolerable provided it does not affect the identification of the direction of thermal expansion and contraction.

[0070] 5) Roundness or ovality deviation

[0071] Allowable error: drum inner diameter deviation <1%, weld edge angle ≤4mm; if exceeded, the design needs to be optimized, but not immediately rejected.

[0072] The review report includes the location of the problem point, problem description, violated standard clauses and rectification suggestions.

[0073] The intelligent identification method proposed in this paper utilizes computer technology and artificial intelligence algorithms to achieve objective and unified identification standards. It is unaffected by subjective factors and consistently conducts identification in accordance with pre-set boiler design standards and inspection regulations, significantly improving the accuracy and reliability of the identification results.

[0074] This multimodal approach enables comprehensive and integrated analysis of boiler design documents. For example, when evaluating a boiler component, not only can the material and performance requirements be understood from the textual annotations, but graphical information can also be used to verify whether the actual structure meets the design requirements, thereby more accurately identifying key points.

[0075] Different types of boilers vary greatly in structure, function, application scenarios, etc., which makes it difficult to accurately capture and identify the key identification points on different boiler design documents. Existing technologies often lack effective methods to deal with this diversity. Although this application faces this difficulty, by collecting a large number of existing drawings as data sets for training, the intelligent identification system can learn the characteristics and laws of different boilers. For example, for power plant boilers and industrial boilers, the system can determine their respective key identification points based on the learning results of the data set. For example, power plant boilers pay more attention to indicators related to thermal efficiency and safety, while industrial boilers may pay more attention to parameters that match the production process.

[0076] While some intelligent appraisal systems exist in other fields, such as architectural drawing appraisal systems and mechanical parts drawing appraisal systems, boiler design requires unique expertise and regulatory requirements. Boiler design involves special requirements such as high temperature, high pressure, and safety, resulting in extremely stringent and detailed design standards and inspection regulations.

[0077] The intelligent appraisal system described in this invention addresses the specific characteristics of boiler design by integrating specialized knowledge and specifications into a multimodal drawing review model. For example, when identifying material information, the system determines whether the selected materials meet relevant standards based on the boiler's operating environment and pressure requirements. When evaluating structural designs, it considers factors such as thermal expansion and stress distribution to ensure design safety and reliability.

[0078] The intelligent identification system for boiler design documents based on multimodal knowledge graph includes:

[0079] A data acquisition module, configured to acquire a large number of existing boiler design drawings as a data set, wherein the data set includes design drawings of various types of boilers;

[0080] The multimodal feature extraction module includes an image feature extraction unit, a text feature extraction unit, and a multimodal fusion unit. The image feature extraction unit is used to extract image information from the boiler design file, and the text feature extraction unit is used to extract text information from the boiler design file. The multimodal fusion unit fuses the extracted image information and text information to form a text-image multimodal joint feature vector.

[0081] The knowledge graph construction module maps the text-image multimodal joint feature vector into entity nodes in a heterogeneous graph, establishes semantic connections with the structural constraint information extracted from the boiler design standard specifications, and constructs a multimodal design knowledge graph with engineering semantics, graph structure association, and reasoning capabilities.

[0082] The intelligent comparison and decision-making module includes an intelligent comparison unit and a decision-making unit. The intelligent comparison unit is used to compare the key review points on the extracted boiler design drawings to be reviewed with the boiler design knowledge graph, and uses the rule engine Drools to perform logical reasoning. The decision-making unit is used to determine whether the drawing design complies with regulations. For problematic design points, the module points out the problems and explains the reasons, and generates a detailed review report.

[0083] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the above method.

[0084] A computer-readable storage medium is used to store a computer program. When the computer program is run on a computer, the computer is caused to execute the above method.

[0085] Compared with the prior art, the present invention has the following beneficial effects:

[0086] 1. Improve identification efficiency:

[0087] The current manual evaluation of boiler design documents is inefficient. A single complex drawing can take an appraiser hours or even days. However, the intelligent evaluation system can rapidly analyze and process drawings in a fraction of the time. For example, a drawing that would have taken a day to evaluate manually can be completed in just minutes by the intelligent evaluation system, significantly shortening the evaluation cycle and improving work efficiency.

[0088] 2. Reduce the missed review rate

[0089] Manual appraisals are prone to omissions due to factors such as negligence. Based on pre-set rules and algorithms, the intelligent appraisal system comprehensively and meticulously evaluates drawings, ensuring that no critical points are missed. For example, when evaluating the weld design of a boiler, manual inspections may overlook subtle issues due to fatigue or negligence. However, the intelligent appraisal system can accurately identify these issues through precise pattern recognition and rule matching, significantly reducing the omission rate.

[0090] 3. Save labor costs

[0091] Traditional manual appraisals require a large number of professional appraisers, resulting in high labor costs. The introduction of intelligent appraisal systems can reduce reliance on human appraisers, lowering the manpower investment required by companies or relevant departments. For example, an appraisal that previously required an entire team may now require only a small number of individuals to simply review the results of the intelligent appraisal system, saving significant labor costs.

[0092] 4. Targeted processing of multimodal data

[0093] Drawing information in different fields has different characteristics and needs to be processed in different ways. In boiler design documents, the text information contains a large number of professional terms and parameters, while the graphic information involves complex piping layouts, pressure vessel structures, etc.

[0094] This application's multimodal drawing review model addresses these characteristics of boiler design documents by processing both textual and graphical information in a targeted manner. For example, when extracting textual information, specialized natural language processing techniques are employed to accurately identify and understand boiler terminology; when processing graphical information, image processing and computer vision techniques are employed to precisely measure and analyze complex structures and dimensions.

[0095] 5. Promote the intelligent development of boiler design and appraisal

[0096] Currently, the level of intelligence in boiler design and appraisal is relatively low. This intelligent appraisal system introduces new technical means and methods to this field. Its application will promote the transformation of boiler design and appraisal work from traditional manual methods to intelligent methods, improving the level and quality of appraisals across the industry.

[0097] 6. Improve the safety and reliability of boiler design

[0098] By accurately identifying and verifying key points in boiler design documents, the intelligent verification system can promptly identify design issues and hidden dangers, ensuring that the boiler design complies with relevant standards and specifications. This helps improve boiler safety and reliability, reduces safety incidents and quality issues caused by design flaws, and provides a strong guarantee for safe boiler operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 This is a flow chart of the boiler design document intelligent identification method based on multimodal knowledge graph according to the present invention;

[0100] Figure 2 This is a flow chart of the multimodal feature extraction process of the present invention;

[0101] Figure 3 This is the HED model architecture diagram for image feature extraction from boiler design drawings according to the present invention;

[0102] Figure 4 This is a diagram of the improved U-Net model architecture of the present invention;

[0103] Figure 5 This is a flowchart of the improved multimodal feature fusion method according to the present invention;

[0104] Figure 6 A flowchart of constructing a multimodal boiler design knowledge graph according to the present invention;

[0105] Figure 7 Schematic diagram of the boiler design document intelligent identification system based on multimodal knowledge graph described in the present invention. DETAILED DESCRIPTION

[0106] The technical solution of the present invention is described in detail below with reference to the embodiments, but the protection scope of the present invention is not limited to the embodiments.

[0107] Example 1:

[0108] like Figure 1 As shown in the figure, the intelligent identification method of boiler design documents based on multimodal knowledge graph includes the following steps:

[0109] Step S1: A data collection module is used to collect a large number of existing boiler design documents as a data set. Boiler design documents include but are not limited to boiler design drawings, technical documents, and 3D model data. Boiler design drawings are acquired using a high-resolution scanner with a resolution greater than 600 dpi; technical documents are acquired from electronic or paper documents using optical character recognition (OCR); and 3D model data is generated using professional 3D modeling software or acquired using 3D laser scanning equipment.

[0110] Step S2: Use the multimodal feature extraction module to extract the information on the boiler design drawings in the data set to form a text-image multimodal joint feature vector. The multimodal feature extraction module includes an image feature extraction unit and a text feature extraction unit.

[0111] like Figure 2 As shown in Figure 2, the multimodal feature extraction method specifically includes:

[0112] Step S2.1: Image Feature Extraction: First, the HED model and Canny edge detection algorithm are used to extract overall edge information from the design drawing, preliminarily determining the drawing's scope and boiler structure. Next, a U-Net model is used to identify the boiler's contour. By adding an attention mechanism, the model's focus on key boiler components (the boiler body, pipes, valves, etc.) is enhanced, improving the accuracy of boiler contour extraction. The attention mechanism dynamically adjusts the model's attention weights for different regions, ensuring that key component features are not overlooked. Each extracted graphic component is then encoded as a high-dimensional image feature vector, retaining its geometric position information and structural feature parameters, such as circles, ellipses, connecting lines, and dimension lines, to facilitate positional alignment and semantic matching with text features.

[0113] In this embodiment, if Figure 3 As shown in the figure, edge detection is specifically as follows: the boiler design drawing image is used as the input of the HED model, and its five convolutional layers (Conv1 to Conv5) and pooling layers gradually extract features. The side-output branch of each convolutional layer generates an edge prediction map, and the resolution is restored by the 1x1 convolution layer and the upsampling layer. The fusion layer FusionLayer performs weighted fusion and Sigmoid activation to generate the final edge map FinalEdgeMap to obtain complex edge information; at the same time, the Canny algorithm is used to perform Gaussian filtering, gradient calculation, non-maximum suppression and double threshold processing on the image to extract regular edges. The combination of the two provides high-quality edge data for subsequent contour recognition.

[0114] In this embodiment, the HED model accurately captures key edge information, such as the boiler body contour: the irregular contour lines of the boiler body; pipe joints: the connections between pipes and components such as the boiler body, valves, and flanges; valve and flange edges: the complex structural edges of valves and flanges; and welds and bolt holes: tiny but important edge details. Compared to traditional edge detection algorithms, the HED model utilizes multi-scale feature fusion technology to simultaneously process both high-resolution and low-resolution edge information, thereby extracting more refined edge structures in boiler design drawings. For example, the HED model accurately captures details such as the connection between the boiler body and pipes, the valve contour, and the flange edge. This capability is particularly important for boiler design drawings, as these details directly determine the functionality and safety of the boiler. While the HED model excels at extracting complex edges, its high model complexity can introduce unnecessary noise when processing simple and continuous edges. To address this, this application incorporates the Canny edge detection algorithm, using steps such as Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding to extract clear and continuous edge lines. In boiler design drawings, the Canny algorithm is particularly suitable for extracting regular edges, such as straight pipes and circular flanges. By combining the HED model with the Canny algorithm, we can effectively reduce noise interference while maintaining edge detection accuracy, ensuring that the extracted edge information is both comprehensive and accurate.

[0115] like Figure 4 As shown in the figure, contour recognition specifically uses an improved U-Net model. Its encoder extracts features through multiple layers of convolution and downsampling, while the decoder restores the feature map resolution through upsampling and convolution. An attention gating module is added to the skip connections to enhance the model's ability to focus on key components. The bottleneck layer captures the highest-level feature information, ensuring that the model simultaneously processes large-scale structures and small-scale details.

[0116] The U-Net model performs contour recognition, and by adding an attention mechanism, it significantly improves the model's ability to pay attention to key components (boiler body, pipes, valves, etc.). With its unique encoder-decoder structure and jump connection, the U-Net model can efficiently capture multi-level feature information in the image, and is particularly suitable for extracting complex contours. On this basis, this application introduces an attention gating module in the jump connection to dynamically adjust the weight distribution of the feature map, so that the model focuses more on the contour area of the key components of the boiler, while suppressing the interference of background noise. In addition, through multi-scale feature fusion technology, the model can simultaneously process large-scale structures (boiler body) and small-scale details (welds, bolt holes), improving the comprehensiveness and accuracy of contour extraction.

[0117] Step S2.2: Text feature extraction: Based on OCR technology, text areas are detected from the boiler design drawings and converted into editable text data. The text information in the drawings is accurately extracted, including key content such as boiler component names, material types, and dimension annotations. Then, using the improved BERT model, the text information extracted by OCR is annotated and encoded into high-dimensional semantic features to form a text feature vector. At the same time, the spatial position information of the text in the drawing is retained as the text input for multimodal fusion.

[0118] Specifically, the optimized OCR technology is used to perform image processing and pattern recognition on boiler design drawings, detect text areas and convert them into editable text data, accurately extract key information such as component names, material types, dimension markings, and effectively handle complex backgrounds, tilted text, and blurred characters, thereby improving the accuracy and robustness of text recognition.

[0119] The advantages of the BERT model lie in its good semantic expression capabilities and powerful transfer learning characteristics. It is suitable for scenarios where professional field corpus is relatively scarce, such as boiler design drawings. It can quickly adapt to the terminology distribution and expression style in industrial drawings through the driving force of terminology vocabulary and industry background knowledge.

[0120] Step S2.3: Multimodal fusion: Figure 5 As shown, first, the image feature vectors and text feature vectors are mapped to the same semantic space and aligned. For example, the contour features of a boiler body are aligned with the corresponding text features, such as component names and dimension annotations, to ensure consistency of multimodal information. Then, the aligned image and text feature vectors are sequentially concatenated to form a multimodal feature sequence. For example, the contour feature vector of the boiler body is sequentially concatenated with the text feature vectors of the corresponding component names and dimension annotations to form a multimodal feature sequence. The self-attention mechanism of the Transformer multimodal fusion module is then used to calculate the correlation weights between feature vectors and dynamically adjust the feature representation. For example, the contour features of the boiler body interact with the text features of the component names and dimension annotations, and the correlation weights between them are calculated, thereby enhancing the semantic information related to the boiler body. A multi-layer Transformer encoder then performs nonlinear transformations and interactions on the multimodal feature sequence, extracting higher-level semantic representations layer by layer, understanding the structural relationships between components, and fusing multimodal features to generate a unified semantic representation that combines image and text information. This accurately describes the semantic content of the boiler design drawings, thereby achieving a comprehensive understanding of the boiler design drawings. This fusion method not only makes up for the shortcomings of single modality information, but also further improves the comprehensiveness and accuracy of feature extraction through the complementarity between features. Specifically, the image feature vector and the text feature vector are mapped to the same semantic space and aligned and spliced, including:

[0121] Step S2.3.1: Perform normalized spatial embedding processing based on the spatial position information of each text annotation and graphic component in the drawing.

[0122] Specifically, the BERT series of models was used. These models, which utilize a bidirectional Transformer architecture to deeply model context, can fully capture the semantic information of field terms such as material types, structural names, and technical parameters within the text on drawings. The BERT model encodes each text annotation into a 768-dimensional semantic feature vector while preserving the text's spatial position (x, y, w, h) within the drawing for subsequent alignment.

[0123] For vectorized processing of graphic information, the system uses the ResNet residual network architecture within a convolutional neural network to encode structural elements in the drawings. This model has strong image feature extraction capabilities and can identify common graphical features found in boiler structures, such as edge contours, dimensioning, and connectors. To accurately extract discriminative local structural regions within the drawings, the ROIAlign algorithm is introduced. This algorithm precisely aligns candidate regions within the feature map based on their locations, avoiding spatial information errors that may occur during image feature pooling.

[0124] Each extracted graphic component is encoded as a 512-dimensional image feature vector, and its geometric position information and structural feature parameters (such as circles, ellipses, connecting lines, dimension lines, etc.) are retained to facilitate position alignment and semantic matching with text features.

[0125] ResNet was chosen because it balances extraction efficiency and expression capabilities, adapting to the complex structural layout of multi-scale and multi-type primitives in boiler design drawings. ROIAlign further improves the accuracy of expressing local features of structural primitives, making it particularly suitable for analyzing boiler components with subtle annotation differences and size sensitivity.

[0126] Subsequently, after completing the preliminary encoding of text and image information, the system needs to align the spatial position relationship between the two in order to achieve cross-modal fusion. At this stage, normalized spatial embedding processing is first performed based on the spatial position information (x, y, w, h) of each text annotation and graphic component in the drawing. Specifically, the system uniformly maps all drawing coordinates to the [0, 1] interval to generate a 128-dimensional position encoding vector, which not only contains the center point coordinates, but also encodes spatial structural information such as the shape, area ratio, and directionality of the primitive, thereby providing a high-precision position information basis for image and text alignment.

[0127] Step S2.3.2: Use a nearest neighbor search method to screen all image-text pairs based on the distance or overlap between spatial location codes, identifying image-text combinations that are physically proximal on the drawing and logically semantically related. Specifically, a K-nearest neighbor algorithm is used to identify a set of candidate image-text matching pairs. This method can quickly identify image-text combinations that are physically proximal on the drawing and potentially semantically related, laying the foundation for subsequent semantic fusion.

[0128] For these candidate image-text pairs, we utilize hint engineering technology to construct refined hint templates, guiding the large language model to analyze whether there are semantic connections between the image and text. By employing a "setting context + pre-set intent" approach, hint engineering allows the large model to leverage its cross-semantic understanding and reasoning capabilities, automatically identifying potential connections between the image and text. For example, for each candidate image-text pair, we generate a hint template as shown in Table 1 below.

[0129] Table 1 Prompt word template

[0130]

[0131] After receiving the prompt, the large language model can automatically perform semantic reasoning, for example, output: "Yes. 'Q345R' in the text is the name of the material of a commonly used boiler pressure component, and 'thickness 20mm' is consistent with the dimension annotation in the drawing, indicating that the text is a material description of the drawing." Through this structured prompt design, the system can greatly improve the accuracy of semantic matching between images and text, and is particularly suitable for industrial design drawing scenarios with dense professional terms and highly diverse graphic and text expressions. Compared with traditional keyword matching or vector similarity calculation, the intelligent matching method based on large models can handle non-explicit image-text relationships, that is, the text does not directly describe the shape of the graphic, but contains structural / material reasoning information; in addition, the intelligent method is interpretable, and the large model can output the reasons to facilitate subsequent comparison or manual review; finally, based on prompt engineering technology, the prompt template can also be updated to support more structural categories and industry terms.

[0132] Step S2.3.3: Retain the image-text matching pairs with high semantic correlation to construct the fused text-image joint feature vector, which is specifically expressed as follows:

[0133]

[0134] Where: V text Represents the text semantic vector; V image Represents the graph structure feature vector; V position represents the position encoding vector; Represents vector concatenation.

[0135] In addition, in order to capture the potential cross-modal implicit connections between images and text, the attention mechanism is introduced in the multimodal feature fusion stage. The specific implementation method is: when constructing the joint feature vector V joint Previously, the system introduced a weighted attention layer for the image-text candidate pairs to model the correlation between different tokens in the text semantic vector, such as "material", "thickness", "Q345R", etc., and the local area features of the image.

[0136] For example, the weighted fusion expression is as follows:

[0137]

[0138] Among them, α i Attention weights represent the degree of attention each text token pays to the target graphic area. The system is trained to enhance its perception of specialized terminology (for example, assigning higher weights to "pressure element," "shell," and "weld"). This mechanism enables the system to highlight key semantic points while deemphasizing irrelevant information when fusing features, thereby improving the discriminability of multimodal vectors.

[0139] Step S2.4: Key parts analysis and

[0140] Based on SHAP technology, the contribution of each feature to the model output is quantified, the importance of key parts such as the boiler body, pipelines, and valves is clarified, and a key feature heat map is generated. This visualization method intuitively shows the distribution of key parts in the drawing and their impact on the overall structure, ensuring that all key features are fully captured.

[0141] Through iterative training and testing, model parameters and feature weights are adjusted according to the SHAP analysis results to improve the accuracy and robustness of key part extraction, ensure that the model is adaptable to different types of boiler design drawings, verify the comprehensiveness and accuracy of feature extraction, and improve the reliability of the model in practical applications.

[0142] Step S3: Take the text-image multimodal joint feature vector output in step S2 as input, map it into entity nodes in the heterogeneous graph, and establish semantic connections with the structural constraint information extracted from the boiler design standard specification clauses, so as to construct a multimodal boiler design knowledge graph with engineering semantics, graph structure association and reasoning capabilities. Figure 6 As shown, specifically including:

[0143] Step S3.1: Basic vector fusion and preparation: Summarize the BERT encoded semantic feature vector, graph structure vector, and normalized position encoding vector generated in step S2, and use the attention weighting mechanism to fuse these three types of vectors to obtain a unified embedding representation, providing basic semantic input for graph node modeling.

[0144] Step S3.2: Construct entity node:

[0145] Step S3.2.1: Semantic type determination: A multi-classification model based on XGBoost, with its excellent nonlinear expression and feature selection capabilities, processes each nested structure that integrates text, graphics, and position semantics in the boiler design drawings, performs semantic type determination, and outputs predefined node types such as "boiler components", "structural parameters", "material types", and "drawing annotations".

[0146] Step S3.2.2: Attribute Parsing and Encapsulation: Based on the node type, the corresponding attribute parsing template is called to perform structured encapsulation of the semantic information, spatial parameters, and original extracted metadata in the embedded vector, generating standardized fields such as component type, thickness, unit, location, and drawing number. For example: Component Type = Water Level Gauge, Thickness = 9mm, Unit = mm, Location = [0.34, 0.48], Drawing Number = DRW-001-A1.

[0147] Step S3.2.3: Node storage and annotation: Format the structured nodes according to the JSON-LD semantic modeling specification, write them to the database by calling NebulaGraph's HTTP interface, and complete type annotation and index registration.

[0148] Step S3.3: Generate Semantic Edge Relationships: Semantic edge relationships are generated through a collaborative process of rule-driven and semantic similarity-driven mechanisms. The rule-driven path is based on predefined relationships within the engineering template, such as establishing "used material" edges between material annotations and components, and "design parameter" edges between dimension annotations and components. The semantic similarity-driven path uses Sentence-BERT to encode node content. Through vector matching, entity pairs with cosine similarity exceeding a threshold are selected as potential edge candidates. Then, based on factors such as spatial proximity and drawing annotation distance, semantic edges such as "structural dependency," "functional attribution," and "drawing annotation direction" are determined.

[0149] Step S3.4: Standard clause processing and integration

[0150] Step S3.4.1: Obtaining and Preprocessing Standard Clauses: Using the Boiler Safety Technical Regulations TSG11-2020 and its supporting regulatory documents as data sources, LlamaIndex is used to perform PDF segmentation preprocessing on the standard document. In this example, the OpenAI text-embedding-ada-002 model is used to construct a clause-level vector index library.

[0151] Step S3.4.2: Clause Extraction and Structuring: For the nodes generated in the drawings, construct an embedding query and retrieve the semantically closest clause paragraphs using the LlamaIndex retrieval interface. The retrieved paragraphs are fed into the DeepSeek-LLM7B-Instruct model deployed in the local Ollama container environment. Prompt engineering is used to construct standardized prompts and extract structured triples.

[0152] For example, "Please extract the structured elements of this clause in the format: (subject, relationship, attribute)." For the clause "The lower visible edge of the water level gauge should be at least 25mm higher than the highest water level," the model returns the result: (water level gauge, lower edge higher than the highest water level, ≥25mm).

[0153] Step S3.4.3: Fusion of standard and design nodes: Perform regularization verification, logical consistency check, and unit normalization on the extracted results. Use Cypher query language to inject the newly generated standard nodes and their constraint relationship edge structures into the NebulaGraph graph to achieve semantic connection between design nodes and standard specifications.

[0154] Step S3.5: Compliance reasoning model construction and training

[0155] Step S3.5.1: Model selection and implementation: The relational graph attention network R-GAT is selected as the main inference model, and the model structure is built based on the PyTorchGeometric framework. The model supports edge type awareness, multi-head attention mechanism and node embedding update, and is used to learn the weight distribution of multi-relation paths in heterogeneous graphs.

[0156] Step S3.5.2: Training sample construction: Select the node of the component to be evaluated from the archived boiler drawing dataset as the center, expand the two-hop adjacent nodes (parameters, materials, standard clause nodes, etc.) outward to form a local heterogeneous subgraph, and use the expert review conclusion (compliant / non-compliant) as the classification label to construct a supervised training sample.

[0157] Step S3.5.3: Model training and optimization: During the training process, the cross-entropy loss function is used, and the graph neural network learns the importance of semantic paths through multiple rounds of message passing, and optimizes the node representation under the edge type weighting mechanism.

[0158] After training, the model can be used to assess whether the key node structures in the drawings under review meet standard and regulatory constraints. The constructed subgraphs to be reviewed are input into the model, which outputs node classification results and their corresponding attention path weights. To enhance the interpretability of the results, a graph visualization module is introduced to dynamically display the key paths used in the model's predictions and their semantic logic, providing path-level diagnostic evidence for manual review.

[0159] Step S3.6: Model generalization and robustness improvement: To improve the generalization and robustness of the model, the system supports graph structure perturbation and data enhancement strategies, including node attribute perturbation, graph structure pruning, historical version difference annotation, multi-region sampling and other mechanisms to ensure that the model has stable reasoning capabilities under different boiler types, structural layouts and design styles.

[0160] Step S3.7: Standard Update and Graph Reconstruction: First, using the Scrapy framework's standard automatic acquisition and structured update module, we targeted crawled standard updates, clause revisions, and technical notifications from platforms such as the State Administration for Market Regulation. We used an OCR engine to identify standard documents, combined with HanLP for syntax analysis, field extraction, and keyword matching, extracting clause numbers, constraint logic, and applicable condition structure fields. We then structured the new clauses and wrote them into the NebulaGraph graph. We established "substitution relationship" edges pointing to the old standard clauses, triggering dependency path updates and embedding reconstruction for relevant design nodes. This enabled automated coordination between graph version evolution and model adaptation, ensuring the system was synchronized with the latest industry standards.

[0161] To summarize, in this embodiment, LlamaIndex+NebulaGraph+DeepSeek is used to construct a standard knowledge system that integrates graphics and text, and the R-GAT graph neural network is combined to realize the deep integration and dynamic reasoning of boiler design drawings and standard semantics. Moreover, through the standard crawling and graph incremental reconstruction mechanism, the automatic coordination of standard synchronization, structure update and model adaptation is realized, providing the core technical support for the intelligent identification of boiler drawing design compliance.

[0162] Step S4: The extracted key point information from the boiler design drawings to be reviewed is input into the knowledge graph R-GAT model. The model outputs node classification results and their corresponding attention path weights. This determines whether each design point complies with the boiler design specification, providing path-level diagnostic evidence for manual review. Problems with the boiler design are identified and their causes explained, generating a review report that includes the location of the problem point, a description of the problem, any violated standard clauses, and corrective action recommendations.

[0163] Among them, different judgment conditions are corresponding to different key points extracted from boiler design drawings, and the problem points in boiler design drawings are divided into problem points that are vetoed and problem points that allow a certain error;

[0164] Among them, the issues and criteria for veto include:

[0165] 1) The safety water level marking is incorrect or missing

[0166] Judgment conditions: The lowest safe water level of the boiler is not clearly marked in the design drawing or the lowest water level is lower than the highest fire limit; the lowest water level of the shell boiler should be 100mm higher than the highest fire limit (75mm for inner diameter ≤1500mm).

[0167] 2) Lack of emergency water discharge device

[0168] Judgment conditions: The power station boiler drum is not equipped with an emergency water discharge device, or the discharge pipe is located below the minimum safe water level;

[0169] 3) The door hole design does not meet the maintenance requirements

[0170] Judgment conditions: The number of manholes, handholes, and headholes does not meet the minimum requirements or the arrangement is not convenient for cleaning and maintenance. If the inner diameter of the boiler drum is ≥800mm and no manhole is set, it is a direct rejection item;

[0171] 4) The key T-joint adopts overlapping structure

[0172] Judgment conditions: The T-joint at the flue gas scouring area does not adopt a full-penetration butt joint structure with groove processing, and there is an overlap connection;

[0173] 5) The weld does not meet the minimum distance requirements and there is high stress concentration

[0174] Judgment conditions: The center line spacing between the drum and furnace welds is less than 3 times the thickness of the steel plate and less than 100mm, which is prone to structural defects and is directly judged as unqualified.

[0175] 6) Welding quality and non-destructive testing

[0176] The pass rate of non-destructive testing (X-ray, ultrasonic) of welds must be 100%, and the film spot checks must cover key areas (weld intersections, T-joints).

[0177] Problem points and judgment conditions that allow a certain degree of error include:

[0178] 1) Detail error of water level gauge scale line:

[0179] Allowable error: ±2mm does not affect safety judgment, and the system prompts manual review.

[0180] 2) The manhole and handhole are slightly smaller in size but have a reasonable structure:

[0181] Allowable error: Due to structural limitations, if the height or diameter is smaller than the standard but has the possibility of entry, it is considered a tolerance design and requires manual judgment;

[0182] 3) Pipe hole layout margin error

[0183] Allowable error: The edge distance or hole spacing is slightly smaller than the standard but not overlapping. When the hole diameter is <60mm, the edge distance requirement can be reduced by 5%.

[0184] 4) Expansion indicator marking deviation

[0185] Permissible error: For Class A boilers, an offset of the expansion indicator of less than 10% is tolerable provided it does not affect the identification of the direction of thermal expansion and contraction.

[0186] 5) Roundness or ovality deviation

[0187] Allowable error: drum inner diameter deviation <1%, weld edge angle ≤4mm; if exceeded, the design needs to be optimized, but not immediately rejected.

[0188] Example 2:

[0189] like Figure 7 As shown in the figure, the intelligent identification system for boiler design documents based on multimodal knowledge graph includes:

[0190] A data acquisition module, configured to acquire a large number of existing boiler design drawings as a data set, wherein the data set includes design drawings of various types of boilers;

[0191] The multimodal feature extraction module includes an image feature extraction unit, a text feature extraction unit, and a multimodal fusion unit. The image feature extraction unit is used to extract image information from the boiler design file, and the text feature extraction unit is used to extract text information from the boiler design file. The multimodal fusion unit fuses the extracted image information and text information to form a text-image multimodal joint feature vector.

[0192] The knowledge graph construction module maps the text-image multimodal joint feature vector into entity nodes in a heterogeneous graph, establishes semantic connections with the structural constraint information extracted from the boiler design standard specifications, and constructs a multimodal design knowledge graph with engineering semantics, graph structure association, and reasoning capabilities.

[0193] The intelligent comparison and decision-making module includes an intelligent comparison unit and a decision-making unit. The intelligent comparison unit is used to compare the key review points on the extracted boiler design drawings to be reviewed with the boiler design knowledge graph, and uses the rule engine Drools to perform logical reasoning. The decision-making unit is used to determine whether the drawing design complies with regulations. For problematic design points, the module points out the problems and explains the reasons, and generates a detailed review report.

[0194] Example 3:

[0195] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the above method.

[0196] Example 4:

[0197] A computer-readable storage medium is used to store a computer program. When the computer program is run on a computer, the computer is caused to execute the above method.

[0198] Application examples:

[0199] Scenario: A boiler factory appraises the drawings of a 100t / h power station boiler;

[0200] 1. Scan the boiler design drawings: drawing images, A0 size, 600dpi; 3D point cloud data, density 1000 points / cm 2 .

[0201] 2. Intelligent identification of boiler design documents:

[0202] Text recognition: Extract information such as material 12Cr1MoV, design pressure 10MPa, etc.

[0203] Graphic recognition: The height of the fillet weld of the boiler tube is measured to be 8m, while the standard requirement is ≥10m.

[0204] Three-dimensional verification: Verify the strength of the header support structure (compare finite element analysis results with standards).

[0205] Output: The appraisal report marked three issues, including insufficient weld leg height of the fillet welds 2 and 6 of the pipe joints and flange thickness deviation. The rectification suggestions complied with Article 3.4 of TSG11-2020.

[0206] Comparative experiment:

[0207] Traditional method: Two appraisers spent four hours and found one problem.

[0208] The system of the present invention completed the appraisal in 0.4 hours and found 2 problems, one of which was missed by manual review.

[0209] Identification efficiency increased by 80%, from 2 hours per sheet to 0.4 hours per sheet.

[0210] The missed review rate dropped from 5% to below 0.5%.

[0211] The above is merely an example of the implementation of the present invention in a specific river basin, but the scope of protection of the present invention is not limited thereto. If a person skilled in the art is inspired by the above and designs methods and embodiments similar to the technical solution without inventive means without departing from the inventive purpose of the present invention, they shall fall within the scope of protection of the present invention.

Claims

1. An intelligent identification method for boiler design documents based on multimodal knowledge graph, characterized by: The steps include: Step S1: using a data acquisition module to collect a large number of existing boiler design files as a data set, the boiler design files including but not limited to boiler design drawings, technical document texts, and three-dimensional model data; Step S2: using a multimodal feature extraction module to extract information on the boiler design drawings in the dataset to form a text-image multimodal joint feature vector; Step S3: Using the text-image multimodal joint feature vector output in step S2 as input, it is structured and mapped into entity nodes in a heterogeneous graph. Semantic connections are established with the structural constraint information extracted from the boiler design standard specifications, thereby constructing a multimodal boiler design knowledge graph with engineering semantics, graph structure association, and reasoning capabilities. Step S4: Input the extracted key point information on the boiler design drawing to be reviewed into the knowledge graph model to determine whether each design point meets the requirements of the boiler drawing design specification; Point out the problems in boiler design and explain the reasons, and generate an audit report.

2. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 1 is characterized in that: In step S1, the boiler design drawing image is obtained by a high-resolution scanner, and the technical document text is obtained from an electronic document or a paper document after OCR recognition; the three-dimensional model data is generated by professional three-dimensional modeling software or collected by a three-dimensional laser scanning device.

3. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 2 is characterized in that: In step S2, the multimodal feature extraction module includes an image feature extraction unit and a text feature extraction unit, and the extraction method specifically includes: Step S2.1: Image feature extraction: First, the HED model and Canny edge detection algorithm are used to extract the overall edge information from the boiler design drawing, preliminarily determining the drawing's scope and boiler structure. Next, the U-Net model is used to identify the boiler's contour. By adding an attention mechanism, the model's ability to focus on key boiler components improves the accuracy of boiler contour extraction. Each extracted graphic component is then encoded into a high-dimensional image feature vector, retaining its geometric position information and structural feature parameters. Step S2.2: Text feature extraction: Using optical character recognition (OCR) technology, text regions are detected from boiler design drawings and converted into editable text data. Accurately extract text information from the drawings, including but not limited to boiler component names, material types, and dimensions. Then, using the improved BERT model, the extracted text information is annotated and encoded into high-dimensional semantic features to form a text feature vector. The spatial location information of the text within the drawings is retained, serving as the text input for multimodal fusion. Step S2.3: Multimodal fusion: First, the image feature vector and the text feature vector are mapped to the same semantic space and aligned to ensure the consistency of multimodal information; then, the aligned image feature vector and text feature vector are sequentially spliced to form a multimodal feature sequence; the self-attention mechanism of the Transformer multimodal fusion module is then used to calculate the correlation weights between the feature vectors and dynamically adjust the feature representation; and the multimodal feature sequence is nonlinearly transformed and interacted with through a multi-layer Transformer encoder to extract higher-level semantic representations layer by layer, understand the structural relationship between components, fuse multimodal features, and generate a unified semantic representation containing comprehensive image and text information to accurately describe the semantic content of the boiler design drawings.

4. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 3 is characterized in that: In step S2.3, the image feature vector and the text feature vector are mapped to the same semantic space, aligned, and spliced, specifically including: Step S2.3.1: Perform normalized spatial embedding processing based on the spatial position information of each text annotation and graphic component in the boiler drawing; Step S2.3.2: Use the nearest neighbor search method to screen all image-text pairs based on the distance or overlap between spatial location codes, and select image-text combinations that are physically adjacent on the drawing and logically semantically related; Step S2.3.3: Retain the image-text matching pairs with high semantic correlation to construct the fused text-image joint feature vector, which is specifically expressed as follows: V joint =[V text ⊕V image ⊕V position ] Where: V text Represents the text semantic vector; V image Represents the graph structure feature vector; V position represents the position encoding vector; ⊕ represents vector concatenation.

5. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 4 is characterized in that: The step S3 of constructing a dynamic knowledge graph specifically includes: Step S3.1: Basis vector fusion and preparation: Summarize the BERT-encoded semantic feature vector, graph structure vector, and normalized position encoding vector generated in step S2, and fuse these three types of vectors using the attention weighting mechanism to obtain a unified embedding representation; Step S3.2: Construct entity node: Step S3.2.1: Semantic type determination: Based on the XGBoost multi-classification model, each nested structure in the boiler design drawing that integrates text, graphics, and positional semantics is processed to determine the semantic type and output the predefined node types of "boiler component", "structural parameter", "material type", and "drawing annotation". Step S3.2.2: Attribute parsing and encapsulation: Call the corresponding attribute parsing template based on the node category, perform structured encapsulation of the semantic information, spatial parameters and original extracted metadata in the embedding vector, and generate standardized fields; Step S3.2.3: Node storage and annotation: Format the structured nodes according to the JSON-LD semantic modeling specification, write them to the database by calling NebulaGraph's HTTP interface, and complete type annotation and index registration; Step S3.3: Generate semantic edge relationships: Semantic edge relationships are generated collaboratively using both rule-driven and semantic similarity-driven mechanisms. The rule-driven path predefines relationships based on project templates, while the semantic similarity-driven path uses Sentence-BERT to encode node content. Entity pairs with cosine similarity exceeding a threshold are selected as potential edge candidates through vector matching. Spatial proximity and drawing annotation distance are then combined to determine whether to establish semantic edges for "structural dependency," "functional attribution," or "drawing annotation pointing." Step S3.4: Standard clause processing and integration: Step S3.4.1: Acquisition and preprocessing of standard clauses: Using the Boiler Safety Technical Regulations TSG11-2020 and its supporting regulatory documents as data sources, LlamaIndex is used to perform PDF segmentation preprocessing on the standard documents, and an AI model is used to construct a clause-level vector index library; Step S3.4.2: Clause Extraction and Structuring: For the nodes generated in the drawing, construct an embedding query and retrieve the semantically closest clause paragraphs through the LlamaIndex retrieval interface. Deploy the recalled paragraph input in the local A1 model, and construct a standardized prompt to extract structured triples using prompt engineering. Step S3.4.3: Fusion of Standard and Design Nodes: Perform regularization verification, logical consistency check, and unit normalization on the extracted results. Use Cypher query language to inject the newly generated standard nodes and their constraint relationship edge structures into the NebulaGraph graph to achieve semantic connection between the design nodes and the standard specifications. Step S3.5: Compliance reasoning model construction and training Step S3.5.1: Model selection and implementation: We select the relational graph attention network (R-GAT) as the main inference model, build the model structure based on the PyTorchGeometric framework, and learn the weight distribution of multiple relational paths in heterogeneous graphs. Step S3.5.2: Training sample construction: Select the node of the component to be reviewed from the archived boiler drawing dataset as the center, expand the two-hop adjacent nodes outward to form a local heterogeneous subgraph, and use the expert review conclusions as classification labels to construct supervised training samples; Step S3.5.3: Model training and optimization: During the training process, the cross-entropy loss function is used, and the graph neural network learns the importance of semantic paths through multiple rounds of message passing, and optimizes the node representation under the edge type weighting mechanism.

6. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 5 is characterized in that: Step S3 also includes updating boiler design specifications and graph reconstruction. Specifically, the following steps are performed: First, the Scrapy framework-based standard automatic acquisition and structured update module is used to crawl standard updates, clause revisions, and technical notifications from platforms such as the State Administration for Market Regulation. An OCR engine is used to identify standard documents, and HanLP is combined to perform syntax analysis, field extraction, and keyword matching, extracting clause numbers, constraint logic, and applicable condition structure fields. The new clauses are then structured and written into the NebulaGraph graph, with "substitution relationship" edges pointing to the old standard clauses. This triggers the update and embedding reconstruction of the dependency paths of relevant design nodes, achieving automatic coordination between graph version evolution and model adaptation, ensuring the system is synchronized with the latest industry standards.

7. The method for intelligent identification of boiler design documents based on multimodal knowledge graph according to claim 6 is characterized in that: In step S4, the intelligent comparison unit uses a rule engine to input the constructed subgraph of the drawing to be reviewed into the trained R-GAT model, outputs the node classification results and their corresponding attention path weights, and the decision unit evaluates whether the key node structure in the drawing to be reviewed meets the standard specification constraints; and through the graph visualization module, dynamically displays the key path and its semantic logic used by the R-GAT model in judgment and prediction, providing path-level diagnostic basis for manual review; Different key points extracted from boiler design drawings correspond to different judgment conditions, and the problem points in boiler design drawings are divided into problem points that are vetoed and problem points that allow a certain error.

8. The intelligent identification system for boiler design documents based on multimodal knowledge graph is characterized by: include: A data acquisition module, configured to acquire a large number of existing boiler design drawings as a data set, wherein the data set includes design drawings of various types of boilers; The multimodal feature extraction module includes an image feature extraction unit, a text feature extraction unit, and a multimodal fusion unit. The image feature extraction unit is used to extract image information from the boiler design file, and the text feature extraction unit is used to extract text information from the boiler design file. The multimodal fusion unit fuses the extracted image information and text information to form a text-image multimodal joint feature vector. The knowledge graph construction module maps the text-image multimodal joint feature vector into entity nodes in a heterogeneous graph, establishes semantic connections with the structural constraint information extracted from the boiler design standard specifications, and constructs a multimodal design knowledge graph with engineering semantics, graph structure association, and reasoning capabilities. The intelligent comparison and decision-making module includes an intelligent comparison unit and a decision-making unit. The intelligent comparison unit is used to compare the key review points on the extracted boiler design drawings to be reviewed with the boiler design knowledge graph, and uses the rule engine Drools to perform logical reasoning. The decision-making unit is used to determine whether the drawing design complies with regulations. For problematic design points, the module points out the problems and explains the reasons, and generates a detailed review report.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Power grid drawing intelligent review method and system based on knowledge graph

    CN120833124A

  • A power grid drawing intelligent review method and system based on a knowledge graph

    CN120833124B

  • Industrial drawing similarity detection method based on deep learning

    CN121074913A

  • Failure mode identification method and device of new energy ship power system, computer equipment, storage medium and computer program product

    CN121808704A