Ship piping design assistance method based on knowledge graph

CN120764652BActive Publication Date: 2026-09-29WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510898958.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2026-09-29
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

[0004]本发明所要解决的技术问题在于针对上述存在的问题,提供基于知识图谱的船舶管系设计辅助方法,解决船舶管系设计知识重用率低,设计人员对设计资源利用困难的问题

Benefits of technology

[0028]本申请的有益效果是:1、本申请提供的基于知识图谱的船舶管系设计辅助方法,通过三元组联合标注策略,显式建模实体对间的多重关系,避免关系重叠漏检;考虑在编码层增强模型对长距离依赖的捕捉能力,解决长实体因位置信息衰减导致的边界模糊问题;引入注意力机制引导模型聚焦结构化上下文区域,提升知识关联的准确性;2、通过引入基于图嵌入的检索技术,将知识图谱中的实体和关系映射到低维向量空间,通过计算向量相似度来快速、全面地召回相关上下文信息。同时,还可以考虑对检索信息的扩写与多路召回、综合排序等方法,不断优化检索策略,提高召回的准确性和全面性;3、结合大语言模型的推理能力和知识图谱的结构信息,运用设计提示词工程与思维链技术,保证模型可以充分利用知识图谱中的逻辑关系和约束条件,对生成过程进行引导和约束,确保生成的答案符合船舶设计专业规范和逻辑要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764652B_ABST
    Figure CN120764652B_ABST
Patent Text Reader

Abstract

The application provides a ship piping system design auxiliary method based on a knowledge graph, utilizes a natural language processing technology to process scattered design documents and design criteria of an enterprise, extracts knowledge triple information, further utilizes retrieval augmentation generation (RAG) technology based on the constructed knowledge graph, realizes subjective query of a graph database design information, analyzes and corrects design information picture-text document data corresponding to design drawings and models, finally realizes design requirement query of a design personnel in a design process and automatic auditing of a design drawing model, the ship piping system design knowledge graph is constructed, design knowledge integration and sharing are completed, and design knowledge utilization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of ship piping system design, specifically relating to a knowledge graph-based auxiliary method for ship piping system design. Background Technology

[0002] The booming development of digital technology has injected powerful transformative momentum into the shipbuilding industry, covering the entire lifecycle from design and manufacturing to operation, maintenance, and even management, and has become a key research area of ​​focus for the shipbuilding industry. In the field of ship design, approximately 60% of the work relies on existing experience. To ensure design compliance, designers must have a comprehensive grasp of relevant design knowledge and past experience. However, the utilization of design knowledge resources in the traditional ship design process has many significant drawbacks. On the one hand, knowledge storage is scattered and disordered, retrieval efficiency is low, and a large amount of experience knowledge lacks structured organization, making it susceptible to knowledge loss due to personnel turnover. On the other hand, the different design requirements of various institutions such as shipowners, classification societies, and construction units further exacerbate the complexity of knowledge management. Therefore, it is urgent to solve a series of key problems such as the analysis, organization, expression, storage, and efficient application of design knowledge to break the current knowledge management dilemma. Knowledge graphs, with their unique advantages, provide new ideas and directions for solving these problems.

[0003] In the field of ship design, combining structured design information from knowledge graphs with the reasoning capabilities of large language models has strong application potential for design information retrieval and compliance assessment. In the early stages of intelligent manufacturing, the construction of knowledge ontology or knowledge graphs primarily relied on manual operation or direct conversion from specific data models. This construction model has significant drawbacks: it is inefficient, prone to errors, and faces numerous difficulties in updating and maintenance, making it difficult to effectively meet the actual needs of future industry development. Thanks to advancements in information technologies such as deep learning and optical character recognition (OCR), as well as the continuous enrichment of engineering documents and data resources, automatically learning and extracting knowledge graphs from massive amounts of data has become a reality. Meanwhile, design review still heavily relies on manual verification, resulting in low levels of intelligence and problems such as inconsistent review standards, susceptibility to errors, human manipulation, and high costs. Combining the storage and retrieval advantages of knowledge graphs with the generation and reasoning capabilities of large language models to serve ship designers in assessing design compliance and retrieving design knowledge is of significant value in the intelligent development of the shipbuilding industry. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a knowledge graph-based auxiliary method for ship piping system design, thereby addressing the problems of low knowledge reuse rate in ship piping system design and difficulty for designers in utilizing design resources.

[0005] The embodiments of this application are implemented as follows: This application provides a knowledge graph-based auxiliary method for ship piping system design, characterized by the following steps: Step a, Knowledge Graph Construction: Step a1: Collect relevant data on ship piping system design and construct a data model, i.e., entity and relationship framework; Step a2: Clean the data, label the data using the entity and relation framework, divide the data into triplet format of head entity-relation-tail entity, and divide the training set and test set according to the ratio to train and evaluate the model. Step a3: Based on the pre-trained language model BERT and the structured convolutional self-attention SCSA entity relation joint extraction model, the absolute position information is incorporated into the feature representation in the form of a rotation matrix through the rotation position encoding (ROPE) mechanism, thereby realizing the joint extraction of entities and relations from the input design information. Step a4: Set the learning rate to 1e-5, the number of training loops to 50, train the model, and evaluate the model performance using evaluation metrics. Step a5: Use the trained model to extract triplet data information and perform verification and cleaning. Import the cleaned entities, relations and attributes into the graph database. Step a6: Identify and merge entities from different data sources, different extraction batches, or different locations in the same data source that represent the same real-world object by calculating similarity, and construct a unified knowledge view; Step b, retrieval filtering and inference generation: Step b1: Extract semantic features from the input query information through word segmentation, padding, truncation preprocessing, average pooling, and L2 normalization. Step b2 introduces graph embedding technology to map entities and relationships in the knowledge graph to a low-dimensional vector space for storage; Step b3: Achieve fast knowledge graph data retrieval based on vector similarity. The query text is converted into a high-dimensional semantic vector through the GTEembedding model, and L2 distance is calculated using the vector index data built by the FAISS library to achieve millisecond-level retrieval. Step b4: The key design information in the query is automatically parsed by the query input key information extraction module to construct structured search conditions and realize direct query matching of the graph database; Step b5: For the data information jointly stored by Neo4j graph structure and FAISS vector space, dual-path retrieval of vectors and keywords is achieved through vector retrieval and graph retrieval, and the retrieved information is comprehensively ranked by semantic similarity. Step b6 involves dynamically constructing prompts to combine user questions with the search context, guiding the model to generate accurate answers.

[0006] In some alternative implementations, step a3 includes the following specific details: First, the input text is encoded using a pre-trained BERT model. BERT captures deep semantic information and contextual dependencies in the text through a multi-layer Transformer structure, generating high-quality text representations. Then, a Structured Convolutional Self-Attention (SCSA) module is introduced into the BERT output. Combining the local feature extraction capability of convolution operations with the global modeling capability of self-attention, four different scale convolutional kernels [3, 5, 7, 9] are used to capture multi-level contextual information, resulting in an attention-enhanced feature matrix X that integrates multi-scale local features. atnn As shown in the following formula:

[0007] Where ⊙ represents element-wise multiplication, X i Grouping by features, k i ∈{3, 5, 7, 9} represents the kernel size; The input features are processed by multi-scale convolution to extract local and global features. After feature concatenation, spatial attention weights are generated by activation function and fused with the original features by dot product to obtain the feature enhancement matrix.

[0008] The self-attention and cluster weighted calculation process of the SCSA module is achieved through the cluster weight Ω. c By introducing prior grouping knowledge, the ability to capture structured features such as entity relationships is improved. The specific formula is as follows:

[0009] Where Q, K, and V are generated by linear transformation of the input sequence features, achieving weighted aggregation of semantic association information in the input sequence, and Ω c Let d be the cluster attention weight matrix. k This is the head dimension; cluster weighting can enhance the model's ability to model groups of different semantic features. Additionally, the expression for the rotation position encoding (ROPE) mechanism mentioned in the steps is as follows:

[0010] Where x represents the encoded sequence information, m represents the sequence position information, and the absolute position information is incorporated into the feature representation through complex rotation while preserving the relative positional relationship, where θ m Let m be the rotation angle between position m and dimension i.

[0011] In some optional implementations, the evaluation metrics described in step a4 include: precision, recall, and F1 score, which are calculated using the following formulas:

[0012]

[0013]

[0014] Where, correct_num is the number of correctly classified entity relation triples in the model, predict_num is the total number of triples identified and extracted by the model, and gold_num is the total number of labeled triples in the dataset.

[0015] In some optional implementations, the triplet data information verification and cleaning described in step a5 includes: deleting erroneously extracted information content and supplementing unidentified extracted triplet information.

[0016] In some alternative implementations, step a6 includes the following specific details: An index is created for the key identifying attributes of entities. Jaccard similarity and Cosine similarity are used to calculate the probability score of the generated candidate entity pairs representing the same object. The calculation formula is as follows:

[0017]

[0018] Where S j For Jaccard similarity, S c Let be the Cosine similarity, (e1, e2) be candidate entity pairs, A(e) be the set of key attribute values ​​of entity e, and v(e) be the vector of key attribute values ​​of entity e. The similarity scores are all in the range of [0, 1], and the larger the value, the more similar the entity. The calculated multiple similarity scores are weighted and summed to form a comprehensive similarity score S. composite :

[0019] Further, a similarity threshold is set, and the candidate pairs are judged to represent the same entity based on the comprehensive similarity score. Multiple representations that are determined to be the same entity are merged into a unified entity node. Then, the alignment results are verified by manual sampling, errors are corrected, and the alignment model is optimized.

[0020] In some alternative implementations, step b1 includes the following specific details: The input query information is preprocessed through word segmentation, padding, and truncation, and semantic features are extracted using average pooling and L2 normalization. The specific formula is as follows:

[0021]

[0022] Where e represents the sentence embedding after average pooling, H represents the embedding matrix of the input token, M is the attention mask, where M[i]=1 indicates a valid token, ε is the minimum value to avoid division by zero, and V norm This represents the normalized vector value, where v represents a vector input. L2 normalization aims to scale the vector to a unit length, eliminating the impact of length differences. Further, based on the large model API, the identification of entities and relationships in the field of ship piping system design is realized through query input, including: dynamically obtaining entity types and relationship types by loading piping system design diagram pattern files to construct a domain-specific extraction range; adopting prompting engineering technology and combining customized JSON format instructions to guide the model to return structured results to ensure output standardization; and controlling the determinism of model output by using the low temperature parameter temperature=0.0 to improve the stability of extraction results.

[0023] In some optional implementations, step b2 includes the following specific content: vectorizing entities and triples in the graph database using a pre-trained language model, constructing a vector index using the FAISS library, and realizing batch processing and persistent index storage; achieving model replaceability through dependency injection to adapt to different embedding models; automatically adapting to CPU / GPU devices to improve the deployment capability of different devices and improving the efficiency of large-scale data processing by combining batch processing mechanisms, ultimately providing efficient semantic vector representation and retrieval capabilities for designing information query tasks.

[0024] In some alternative implementations, step b3 includes the following specific details: This system enables rapid retrieval of knowledge graph data based on vector similarity. The query text is converted into a high-dimensional semantic vector using the GTE Embedding model. L2 distance is calculated using vector index data built with the FAISS library, achieving millisecond-level retrieval. The L2 distance calculation expression and conversion formula are as follows:

[0025] Where d is the linear distance between two vectors in Euclidean space, and x and y are two n-dimensional vectors. The smaller the distance, the higher the similarity between the vectors. In vector retrieval, similarity is the similarity score in the interval [0,1] after converting the L2 distance. This conversion makes the similarity 1 when the distance d=0, and the larger the distance, the closer the similarity is to 0, which is convenient for sorting the results. A dynamic mapping mechanism is adopted to convert the search results into structured text containing type, name and description, thereby improving the readability of the results; the number of returned results is controlled by the top_k number parameter, and the Euclidean distance is converted into a similarity score in the [0,1] interval, which facilitates the sorting of results.

[0026] In some optional implementations, step b4 includes the following specific content: using a multi-dimensional retrieval strategy, graph traversal is implemented using the Cypher query language, including direct entity matching, relation pattern matching, and graph pattern association retrieval; fuzzy matching of node attributes and relation type constraints are supported, and adjacent nodes and relationships of entities are automatically associated and displayed; finally, the retrieval results are sorted through a relevance scoring mechanism, and direct matching results are given higher weight.

[0027] In some optional implementations, step b5 includes the following specific content: first, the candidate set is filtered by rules, and then the BERT-based Global Ranking model is used to refine the ranking based on semantic similarity and knowledge graph relevance to ensure a balance between result relevance and efficiency; in addition, the output scale is dynamically controlled by the top_k number parameter to support automatic device detection and provide an error fallback mechanism. When the model fails to load, a simple model reordering is used to enhance the robustness of the module.

[0028] The beneficial effects of this application are as follows: 1. The knowledge graph-based auxiliary method for ship piping system design provided in this application explicitly models multiple relationships between entity pairs through a triplet joint annotation strategy, avoiding missed detections due to overlapping relationships; it considers enhancing the model's ability to capture long-distance dependencies at the encoding layer, solving the boundary ambiguity problem caused by the attenuation of positional information for long entities; and it introduces an attention mechanism to guide the model to focus on structured context regions, improving the accuracy of knowledge association; 2. By introducing graph embedding-based retrieval technology, entities and relationships in the knowledge graph are mapped to a low-dimensional vector space, and relevant contextual information is quickly and comprehensively recalled by calculating vector similarity. Simultaneously, methods such as expanding and multi-path recall of retrieved information and comprehensive ranking can be considered to continuously optimize the retrieval strategy and improve the accuracy and comprehensiveness of recall; 3. Combining the reasoning ability of a large language model and the structural information of the knowledge graph, and using design prompt word engineering and thought chain technology, it ensures that the model can fully utilize the logical relationships and constraints in the knowledge graph to guide and constrain the generation process, ensuring that the generated answers conform to ship design professional specifications and logical requirements. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating the process of constructing the atlas in the embodiments of this application; Figure 2 This is a model structure diagram in the embodiments of this application; Figure 3 This is a flowchart illustrating the question-and-answer implementation process in an embodiment of this application. Figure 4 This is a visualization of the knowledge graph data of the piping system design in the embodiments of this application; Figure 5 This is a diagram illustrating the implementation effect of the piping system design knowledge Q&A in the embodiments of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0032] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0033] It should be understood that the sequence number of each step in the embodiment does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0034] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0035] The features and performance of this application will be further described in detail below with reference to the embodiments.

[0036] The knowledge graph-based auxiliary method for ship piping system design provided in this application includes a knowledge graph construction module, a design knowledge retrieval and filtering module, and a reasoning and generation module.

[0037] like Figure 1 As shown, the knowledge graph construction process includes the following steps: (1) Design of the ontology schema layer: Collect relevant data on ship piping system design and construct a data model, namely the entity and relation framework. The design of the ontology schema layer serves as the structural foundation of the ship piping system design knowledge graph. The entity relation definitions in the data model include: Entity definitions include Attribute, Component, Requirements, Context, Measure, content, Labeling, File, department\people. Relationship definitions include: Traitof, Applyto, Conditionof, Include, Participatein. The specific meanings of entity and relation definitions are shown in Table 1 below.

[0038] Table 1 Entity and Relationship Definitions

[0039] (2) Data labeling: Clean the collected design data, delete irrelevant information, label the data using a predefined entity and relation framework, divide the data into a triplet format of head entity-relation-tail entity, and finally divide the prepared dataset into training set and test set in an 8:2 ratio to train and evaluate the model.

[0040] (3) Model Design: In this invention, a joint entity relation extraction model based on the pre-trained language model BERT and structured convolutional self-attention is proposed. The overall structure of the model is shown in [reference needed]. Figure 2 This model aims to improve the extraction performance of entity relation triples by leveraging the powerful semantic representation capabilities of pre-trained language models and the multi-scale feature extraction capabilities of structured convolutional self-attention mechanisms. Specifically, the model first encodes the input text using a pre-trained BERT model. BERT, through its multi-layered Transformer structure, can capture deep semantic information and contextual dependencies in the text, thereby generating high-quality text representations. Subsequently, a Structured Convolutional Self-Attention (SCSA) module is introduced into the output of BERT. Combining the local feature extraction capabilities of convolutional operations with the global modeling capabilities of self-attention, four different scale convolutional kernels ([3, 5, 7, 9]) are used to capture multi-level contextual information, and the global perception capability of self-attention is integrated to enhance the model's ability to capture entity boundaries and relational features.

[0041] The working process of the structured convolutional self-attention module is as follows: Attention-enhanced feature matrix that integrates multi-scale local features: (1) In the formula: X atnn The features are multi-scale fused, and ⊙ represents element-wise multiplication, where X i Grouping by features, k i ∈{3, 5, 7, 9} represents the kernel size.

[0042] The input features are processed through multi-scale convolution to extract local and global features. After feature concatenation, spatial attention weights are generated by an activation function and then fused with the original features through a dot product. The feature values ​​at each location are reweighted by the attention weights to highlight key semantic regions.

[0043] Module self-attention and cluster weighted computation: (2) Where Q, K, and V are generated by linear transformation of the input sequence features, achieving weighted aggregation of semantic association information in the input sequence, and Ω c Let d be the cluster attention weight matrix. k This is the head dimension; cluster weighting can enhance the model's ability to model groups of different semantic features. To further mine the structured information in the text and improve the accuracy and robustness of entity relation extraction, a Rotation Position Encoding (ROPE) mechanism is used to incorporate absolute position information into the feature representation in the form of a rotation matrix, thus solving the performance degradation problem of traditional position encoding in long sequence tasks. This enhances the model's sensitivity to entity positions and relations; the encoding mechanism expression is shown in Equation (3). Subsequently, the model explicitly constructs all possible entity pair representations, concatenates the head entity representation and tail entity representation, and obtains the joint feature representation of the entity pair through nonlinear projection transformation. Finally, relation classification is performed based on the entity pair features, generating a three-dimensional score matrix of relation-entity pairs. End-to-end training enables the joint extraction of entities and relations from the input design information.

[0044] (3) Where θ m Let x be the rotation angle between position m and dimension i, where x represents the encoded sequence information and m represents the sequence position information. The absolute position information is incorporated into the feature representation through complex rotation, while the relative position relationship is preserved.

[0045] (4) Model Training and Evaluation: The learning rate was set to 1e-5, and the number of training iterations was set to 50. After training, the model performance was evaluated using precision, recall, and F1 score to verify the model's ability to extract ship design knowledge. The above three evaluation indicators were calculated using the following formulas: (4) (5) (6) Where, correct_num is the number of correctly classified entity relation triples in the model, predict_num is the total number of triples identified and extracted by the model, and gold_num is the total number of labeled triples in the dataset.

[0046] (5) Graph verification: Verify the extracted triplet data, delete incorrectly extracted information, and supplement any unidentified extracted triplet information to improve the quality of the graph. Import the cleaned entities, relations, and attributes into the graph database.

[0047] (6) Entity alignment: Identify and merge entities from different data sources, different extraction batches, or different locations in the same data source that represent the same real-world object by calculating similarity, eliminating redundancy and building a unified knowledge view.

[0048] The specific process is as follows: An index is created for the key identifying attributes of entities; Jaccard similarity and Cosine similarity are used to calculate the probability score of the generated candidate entity pairs representing the same object. The calculation formula is as follows: (7) (8) Where S j For Jaccard similarity, S c Let be the Cosine similarity, (e1, e2) be the candidate entity pair, A(e) be the set of key attribute values ​​of entity e, and v(e) be the vector of key attribute values ​​of entity e. The similarity scores are all in the range of [0, 1], and the larger the value, the more similar the entity.

[0049] The calculated multiple similarity scores are weighted and summed to form a comprehensive similarity score S. composite : (9) A similarity threshold is further set, and the candidate pairs are judged to represent the same entity based on the comprehensive similarity score. Multiple representations that are determined to be the same entity are merged into a unified entity node, and then the alignment results are verified by manual sampling to correct errors and optimize the alignment model.

[0050] like Figure 3As shown, by utilizing the constructed knowledge graph related to ship piping system design, the Neo4j database can be queried directly using simple graph database query statements. For complex requirements, the design rules can be queried by matching ternary chained design rules, and the design information related to drawings and models can be reviewed, thereby providing design assistance to designers. The workflow of the design knowledge retrieval and filtering module and the reasoning and generation module is as follows: (1) Key information extraction module for query input: The input query information is preprocessed through word segmentation, padding, truncation, etc., and semantic features are extracted by average pooling and L2 normalization. The specific formula is as follows: (10) (11) Where e represents the sentence embedding after average pooling, H represents the embedding matrix of the input token, M is the attention mask, where M[i]=1, indicating a valid token, and ε is the minimum value to avoid division by zero. V norm This represents the normalized vector value, where v represents a single vector input. L2 normalization aims to scale the vector to a unit length, eliminating the effects of length differences.

[0051] Further, based on the large model API, the system identifies entities and relationships in the field of ship piping system design based on query input. Key technical points include: dynamically obtaining entity and relationship types by loading piping system design diagram pattern files to construct a domain-specific extraction range; employing prompting engineering techniques combined with customized JSON format instructions to guide the model to return structured results, ensuring output standardization; and controlling the determinism of model output through a low-temperature parameter (temperature=0.0) to improve the stability of the extraction results.

[0052] (2) Graph Vectorization Module: Introducing graph embedding technology, entities and relations in the knowledge graph are mapped to a low-dimensional vector space for storage, providing a data foundation for the vector retrieval module. Specifically, a pre-trained language model is used to vectorize entities and triples in the graph database; a vector index is built using the FAISS library to achieve batch processing and persistent index storage; dependency injection enables model replacement, adapting to different embedding models; automatic adaptation to CPU / GPU devices improves deployment capabilities on different devices; and a batch processing mechanism enhances the efficiency of large-scale data processing, ultimately providing efficient semantic vector representation and retrieval capabilities for information query tasks.

[0053] (3) Vector Retrieval Module: Based on vector similarity, it realizes fast retrieval of knowledge graph data. The query text is converted into a high-dimensional semantic vector through the GTE Embedding model. The L2 distance is calculated using the vector index data built by the FAISS library to achieve millisecond-level retrieval. The L2 distance calculation expression and conversion formula are as follows: (12) Where d is the linear distance between two vectors in Euclidean space, and x and y are two n-dimensional vectors, the smaller the distance, the higher the similarity between the vectors. In vector retrieval, similarity is the similarity score in the interval [0,1] after converting the L2 distance. This conversion makes the similarity 1 when the distance d=0, and the larger the distance, the closer the similarity is to 0, which is convenient for sorting the results.

[0054] A dynamic mapping mechanism is adopted to convert the search results into structured text containing type, name and description, thereby improving the readability of the results; the number of returned results is controlled by the top_k number parameter, and the Euclidean distance is converted into a similarity score in the [0,1] interval, which facilitates the sorting of results.

[0055] (4) Neo4j graph retrieval module: The key design information in the query is automatically parsed by the key information extraction module, and the structured search conditions are constructed to realize direct query matching of graph database. The specific implementation method is as follows: the graph traversal is realized by using the Cypher query language through multi-dimensional search strategy, including direct entity matching, relation pattern matching and graph pattern association retrieval; it supports fuzzy matching of node attributes and relation type constraints, and automatically associates and displays adjacent nodes and relations of entities; finally, the search results are sorted through the relevance scoring mechanism, and the direct matching results are given higher weight.

[0056] (5) Information Ranking Module: For data information jointly stored in Neo4j graph structure + FAISS vector space, a dual-path recall of vectors and keywords is achieved through the vector retrieval module and the graph retrieval module to improve the information matching ability, and further sort the recalled information by semantic similarity. The specific implementation method is as follows: firstly, the candidate set is filtered by rules, and then the BERT-based Global Ranking model is used to refine the ranking according to semantic similarity and knowledge graph relevance to ensure a balance between result relevance and efficiency; in addition, the output scale is dynamically controlled by the top_k number parameter to support automatic device detection and provide an error fallback mechanism. When the model fails to load, a simple model re-ranking is adopted to enhance the robustness of the module.

[0057] (6) Generation module: By dynamically constructing prompts (Prompt Engineering), the user's question is combined with the search context to guide the model to generate accurate answers; the results of vector search and graph search are converted into structured knowledge base text using formatting functions, and key conclusions are accompanied by knowledge tracing functions to retain source information and enhance traceability; by strictly limiting the model to answering only based on knowledge base content, the generation of unfounded information is avoided, ensuring the professionalism and reliability of the answers.

[0058] This method uses a knowledge graph of ship piping system design to achieve intelligent question answering, providing designers with an efficient way to query and retrieve design knowledge and reason based on design rules, thus providing strong support for the design process.

[0059] Example 1 To address the issues of long entity names and multiple nested levels in ship design texts, an entity relationship model based on a pre-trained language model and a structured attention mechanism was proposed to fit the extraction model. The evaluation metrics used in the established ship design triplet dataset were precision, recall, and F1 score, and the specific data are shown in Table 2 below.

[0060] Table 2 Calculation Results of Evaluation Indicators

[0061] 1514 data points were marked based on the collected design data (see...). Figure 4 The dataset was divided into training and testing sets in an 8:2 ratio. The number of entity types in the dataset is as follows: Attribute: 764, Component: 1447, Requirements: 586, Context: 631, measure: 1410, content: 45, labeling: 121, File: 258, department\people: 44. The number of relationship types in the dataset is as follows: Traitof: 385, Applyto: 1441, Conditionof: 631, include: 152, participatein: 44.

[0062] Experimental results show that the model achieves an accuracy of 0.7183 and an F1 score of 0.6758 on a relatively limited dataset, demonstrating relatively accurate and usable information extraction. By comparing it with other entity relationship joint extraction models, and considering the feature complexity of the text and the limited dataset, the experimental results demonstrate the advantages and usability of the model proposed in this invention in the construction of end-to-end design knowledge graphs in the field of ship piping system design.

[0063] For the technical support process of this solution in the design flow, by inputting questions or design information diagrams related to ship piping system design, the system's backend rag workflow function will be matched. The input questions will be processed to extract and vectorize entity relationships. For design information diagrams, they will first be parsed using rule templates before entity extraction and vectorization. Further, multi-way matching will be performed on the triplet information, and the matched design information will be sorted and passed as context to the large model. Utilizing the information organization capabilities of the large model and the background information of the ship design drawings, professional responses will be provided to guide the designers' design process. A specific implementation diagram of the piping system design knowledge Q&A is shown below. Figure 5 As shown.

Claims

1. A knowledge graph-based auxiliary method for ship piping system design, characterized in that, Includes the following steps: Step a, Knowledge Graph Construction: Step a1: Collect relevant data on ship piping system design and construct a data model, i.e., entity and relationship framework; Step a2: Clean the data, label the data using the entity and relation framework, divide the data into triplet format of head entity-relation-tail entity, and divide the training set and test set according to the ratio to train and evaluate the model. Step a3, based on the pre-trained language model BERT and the structured convolutional self-attention model SCSA, uses the rotation position encoding (ROPE) mechanism to incorporate absolute position information into the feature representation in the form of a rotation matrix, thereby achieving joint extraction of entities and relationships from the input design information; including the following specific contents: First, the input text is encoded using a pre-trained BERT model. BERT captures deep semantic information and contextual dependencies in the text through a multi-layer Transformer structure, generating high-quality text representations. Then, a Structured Convolutional Self-Attention (SCSA) module is introduced into the BERT output. Combining the local feature extraction capability of convolution operations with the global modeling capability of self-attention, four different scale convolutional kernels [3, 5, 7, 9] are used to capture multi-level contextual information, resulting in an attention-enhanced feature matrix X that integrates multi-scale local features. atnn As shown in the following formula: Where ⊙ represents element-wise multiplication, X i Grouping by features, k i ∈{3, 5, 7, 9} represents the kernel size; the input features are processed by multi-scale convolution to extract local and global features, and after feature concatenation, spatial attention weights are generated by activation function and fused with the original features by dot product to obtain the feature enhancement matrix; The self-attention and cluster weighted calculation process of the SCSA module is as follows, using the cluster weight Ω c Introducing prior grouping knowledge enhances the ability to capture the structured features of entity relationships, as shown in the following formula: Where Q, K, and V are generated by linear transformation of the input sequence features, achieving weighted aggregation of semantic association information in the input sequence, and Ω c Let d be the cluster attention weight matrix. k This is the head dimension; cluster weighting can enhance the model's ability to model groups of different semantic features. Additionally, the expression for the rotation position encoding (ROPE) mechanism mentioned in the steps is as follows: Where x represents the encoded sequence information, m represents the sequence position information, and the absolute position information is incorporated into the feature representation through complex rotation while preserving the relative positional relationship, where θ m Let m be the rotation angle between position m and dimension i; Step a4: Set the learning rate to 1e-5, the number of training loops to 50, train the model, and evaluate the model performance using evaluation metrics. Step a5: Use the trained model to extract triplet data information and perform verification and cleaning. Import the cleaned entities, relations and attributes into the graph database. Step a6: Identify and merge entities from different data sources, different extraction batches, or different locations in the same data source that represent the same real-world object by calculating similarity, and construct a unified knowledge view; Step b, retrieval filtering and inference generation: Step b1 involves preprocessing the input query information through word segmentation, padding, truncation, average pooling, and L2 normalization to extract semantic features; this includes the following specific steps: The input query information is preprocessed through word segmentation, padding, and truncation, and semantic features are extracted using average pooling and L2 normalization. The specific formula is as follows: Where e represents the sentence embedding after average pooling, H represents the embedding matrix of the input token, M is the attention mask, where M[i]=1 indicates a valid token, ε is the minimum value to avoid division by zero, and V norm This represents the normalized vector value, where v represents a vector input. L2 normalization aims to scale the vector to a unit length, eliminating the impact of length differences. Further, based on the large model API, the system identifies entities and relationships in the field of ship piping system design by querying input, including: dynamically obtaining entity and relationship types by loading piping system design diagram pattern files to construct a domain-specific extraction range; employing prompting engineering techniques and customized JSON format instructions to guide the model to return structured results, ensuring output standardization; and controlling the determinism of model output by using a low temperature parameter (temperature=0.0) to improve the stability of extraction results. Step b2 introduces graph embedding technology to map entities and relationships in the knowledge graph to a low-dimensional vector space for storage; Step b3: Achieve fast knowledge graph data retrieval based on vector similarity. The query text is converted into a high-dimensional semantic vector through the GTEembedding model, and L2 distance is calculated using the vector index data built by the FAISS library to achieve millisecond-level retrieval. Step b4: The key design information in the query is automatically parsed by the query input key information extraction module to construct structured search conditions and realize direct query matching of graph database; Step b5: For the data information jointly stored by Neo4j graph structure and FAISS vector space, dual-path retrieval of vectors and keywords is achieved through vector retrieval and graph retrieval, and the retrieved information is comprehensively ranked by semantic similarity. Step b6 involves dynamically constructing prompts to combine user questions with the search context, guiding the model to generate accurate answers.

2. The knowledge graph-based auxiliary method for ship piping system design according to claim 1, characterized in that, The evaluation metrics mentioned in step a4 include: precision, recall, and F1 score. These three metrics are calculated using the following formulas: Where, correct_num is the number of correctly classified entity relation triples in the model, predict_num is the total number of triples identified and extracted by the model, and gold_num is the total number of labeled triples in the dataset.

3. The knowledge graph-based auxiliary method for ship piping system design according to claim 1, characterized in that, The triplet data information verification and cleaning described in step a5 includes: deleting incorrectly extracted information and supplementing unidentified extracted triplet information.

4. The knowledge graph-based auxiliary method for ship piping system design according to claim 2, characterized in that, Step a6 includes the following specific details: An index is created for the key identifying attributes of entities. Jaccard similarity and Cosine similarity are used to calculate the probability score of the generated candidate entity pairs representing the same object. The calculation formula is as follows: Where S j For Jaccard similarity, S c Let be the Cosine similarity, (e1, e2) be candidate entity pairs, A(e) represent the set of key attribute values ​​of entity e, and v(e) represent the vector of key attribute values ​​of entity e. The similarity scores are all in the range of [0,1], and the larger the value, the more similar the entity. The calculated multiple similarity scores are weighted and summed to form a comprehensive similarity score S. composite : Further, a similarity threshold is set, and the candidate pairs are judged to represent the same entity based on the comprehensive similarity score. Multiple representations that are determined to be the same entity are merged into a unified entity node. Then, the alignment results are verified by manual sampling, errors are corrected, and the alignment model is optimized.

5. The knowledge graph-based auxiliary method for ship piping system design according to claim 4, characterized in that, Step b2 includes the following specific steps: using a pre-trained language model to vectorize entities and triples in the graph database, constructing a vector index using the FAISS library to achieve batch processing and persistent index storage; implementing model replaceability through dependency injection to adapt to different embedding models; automatically adapting to CPU / GPU devices to improve deployment capabilities on different devices and combining batch processing mechanisms to improve the efficiency of large-scale data processing, ultimately providing efficient semantic vector representation and retrieval capabilities for designing information query tasks.

6. The knowledge graph-based auxiliary method for ship piping system design according to claim 5, characterized in that, Step b3 includes the following specific details: This system enables rapid retrieval of knowledge graph data based on vector similarity. The query text is converted into a high-dimensional semantic vector using the GTE Embedding model. L2 distance is calculated using vector index data built with the FAISS library, achieving millisecond-level retrieval. The L2 distance calculation expression and conversion formula are as follows: Where d is the linear distance between two vectors in Euclidean space, and x and y are two n-dimensional vectors. The smaller the distance, the higher the similarity between the vectors. In vector retrieval, similarity is the similarity score in the interval [0,1] after converting the L2 distance. This conversion makes the similarity 1 when the distance d=0, and the larger the distance, the closer the similarity is to 0, which is convenient for sorting the results. A dynamic mapping mechanism is adopted to convert the search results into structured text containing type, name and description, thereby improving the readability of the results; the number of returned results is controlled by the top_k number parameter, and the Euclidean distance is converted into a similarity score in the [0,1] interval, which facilitates the sorting of results.

7. The knowledge graph-based auxiliary method for ship piping system design according to claim 6, characterized in that, Step b4 includes the following specific content: using a multi-dimensional retrieval strategy, graph traversal is implemented using the Cypher query language, including direct entity matching, relation pattern matching, and graph pattern association retrieval; fuzzy matching of node attributes and relation type constraints are supported, and adjacent nodes and relationships of entities are automatically associated and displayed; finally, the retrieval results are sorted through a relevance scoring mechanism, and direct matching results are given higher weight.

8. The knowledge graph-based auxiliary method for ship piping system design according to claim 6, characterized in that, Step b5 includes the following specific content: First, the candidate set is filtered by rules, and then the BERT-based Global Ranking model is used to refine the ranking based on semantic similarity and knowledge graph relevance to ensure a balance between result relevance and efficiency; in addition, the output scale is dynamically controlled by the top_k number parameter to support automatic device detection and provide an error fallback mechanism. When the model fails to load, a simple model reordering is used to enhance the robustness of the module.

Citation Information

Patent Citations

  • Knowledge graph construction method and device for ship power system design task

    CN115905574A

  • Ship trajectory prediction method based on dynamic and static knowledge graph joint reasoning

    CN117236495A