EPC field-oriented knowledge sharing platform construction method

By building an EPC knowledge sharing platform based on knowledge graphs and deep learning algorithms, the accuracy and fusion difficulties of the knowledge extraction model in EPC projects are solved, automatic update and security management of the knowledge base are realized, and knowledge management and decision-making efficiency are improved.

CN120598002APending Publication Date: 2025-09-05CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510605639.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing technology has problems in the EPC project, such as insufficient accuracy and comprehensiveness of the knowledge extraction model, difficulty in knowledge integration, limited semantic understanding ability and lack of sensitive information management, resulting in incomplete knowledge base construction, information lag and security risks.

Method used

Based on the knowledge graph, combined with deep learning algorithms and large language models, the EPC knowledge sharing platform is built through graph database pattern layer design, entity recognition, relationship extraction, knowledge fusion and reasoning model to realize automatic update and self-growth of knowledge, and provide intelligent knowledge query and management.

Benefits of technology

It realizes the comprehensive and accurate sharing and utilization of EPC project knowledge, improves knowledge management efficiency, reduces human and material resources investment, and enhances information security and system response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598002A_ABST
    Figure CN120598002A_ABST
Patent Text Reader

Abstract

The invention provides an EPC field-oriented knowledge sharing platform construction method, which is characterized in that a knowledge graph is used as a basis, a deep learning algorithm and an LLM large language model are used as technical supports, and a knowledge extraction model which takes each professional knowledge in the EPC field as a core target is optimized on the basis of the existing knowledge extraction technology and knowledge query technology, so that the knowledge sharing platform construction efficiency is improved. A knowledge fusion module taking an entity alignment technology as a core is added, a knowledge automatic reasoning and complementing module is added, and a sensitive information processing module is added. A professional knowledge question-answering system is constructed through a general large language model and a knowledge graph, and a promote project is used for processing, so that the knowledge reasoning ability is improved, and deep knowledge mining is realized. According to the method, internal knowledge assets in the EPC engineering field can be better shared and utilized, manpower and material resource investment caused by knowledge management is reduced, and cost reduction and benefit increase are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of EPC engineering project knowledge sharing, and in particular to a method for constructing a knowledge sharing platform oriented to the EPC field. Background Art

[0002] EPC projects involve multiple organizations, departments, and disciplines. The knowledge generated during the project process is often underutilized and difficult to manage uniformly. Effectively preserving and passing on the experience and knowledge accumulated in EPC projects has become a key means for all parties to enhance their innovation capabilities and competitiveness. As a means of expressing knowledge in an explicit and organized way, building an EPC knowledge sharing platform can effectively collect, refine, and transfer knowledge, unlock the value of implicit knowledge, improve work efficiency, and promote knowledge utilization and innovation.

[0003] There are various approaches to building a knowledge sharing platform. One approach uses documents as a medium, utilizing ETL (Extract, Transform, Load) tools to classify, organize, and archive diverse knowledge resources, storing them uniformly in a single data resource pool, thereby building an enterprise-level document information management system. Furthermore, document retrieval can be used to establish expert databases and knowledge maps, further enabling efficient knowledge search and knowledge Q&A, thereby improving knowledge management within the enterprise. Another approach uses knowledge graphs as a medium, leveraging advanced knowledge extraction models to store diverse knowledge in the form of triples (subject-verb-object) in a graph database. This approach not only enables structured knowledge processing during data input but also enables dynamic knowledge updating and natural growth through machine learning algorithms and automated rules. Furthermore, the powerful query capabilities of graph databases enable deep-seated knowledge relationships to be discovered, enabling in-depth knowledge mining and intelligent analysis. This approach can also be combined with natural language processing technology to convert user queries into Cypher statements with approximate vectors, enabling more precise knowledge queries and intelligent knowledge Q&A, significantly improving the accuracy and efficiency of knowledge retrieval. This approach is particularly suitable for scenarios that require processing large amounts of complex data and relationships. It helps EPC project participants achieve more efficient utilization and dissemination of knowledge management, while tapping into the potential value of knowledge and helping companies gain a competitive advantage.

[0004] Current shortcomings of existing technologies include: When dealing with the diverse and complex knowledge involved in EPC projects, existing knowledge extraction models often lack accuracy and comprehensiveness. When processing large amounts of heterogeneous data, they may not accurately extract all key knowledge points, resulting in a suboptimal knowledge structure and incomplete knowledge base construction. Furthermore, since most knowledge extraction models lack self-updating capabilities, knowledge bases cannot automatically update as projects progress and new knowledge is generated, causing information lag and incompleteness. EPC projects involve the coordination and collaboration of multiple disciplines. The disparity in knowledge between disciplines and the diversity of data formats make knowledge fusion extremely difficult. Existing models are prone to data conflicts, redundancy, or loss when handling complex knowledge fusion tasks, resulting in poor knowledge base integration. Traditional natural language processing (NLP) technology, when applied to knowledge sharing platforms, typically uses rule functions to approximate user queries into Cypher statements, allowing queries to be performed within a graph database. While this approach can retrieve knowledge, it relies on existing knowledge in the database. If the query involves knowledge points that are not yet recorded or connected in the database, the system cannot provide an effective answer. Furthermore, this processing approach, based on templates and predefined rules, lacks reasoning capabilities and is difficult to handle complex semantic understanding and deep knowledge mining. It can only mechanically return content from the database, unable to provide deeper analysis or recommendations based on context or related knowledge. Existing knowledge sharing platforms often neglect the meticulous management and protection of sensitive information, potentially exposing ordinary users to inadvertent access to sensitive information that should not be made public. This not only poses information security risks but can also lead to a decline in the trustworthiness of knowledge sharing platforms. Summary of the Invention

[0005] In response to the problems of the above-mentioned existing technical methods, such as insufficient accuracy and comprehensiveness, difficulty in knowledge integration, limited semantic understanding ability and lack of sensitive information management, the present invention provides a method for constructing a knowledge sharing platform for the EPC field based on knowledge graph and technically supported by deep learning algorithms and LLM large language models. This method can better share and utilize internal knowledge assets in the EPC engineering field, reduce the human and material resource investment caused by knowledge management, and achieve cost reduction and efficiency improvement.

[0006] In order to achieve the above technical features, the purpose of the present invention is to achieve the following: a method for constructing a knowledge sharing platform for the EPC field, the method comprising:

[0007] Based on the knowledge types in EPC engineering management, carry out graph database model layer design and data collection;

[0008] Based on the entity and relationship information in EPC engineering management, train specialized entity recognition and relationship extraction models;

[0009] Design and train the EPC knowledge fusion model to perform entity alignment operations;

[0010] Design and train knowledge reasoning models to perform knowledge completion operations;

[0011] Build a question-answering system framework based on large models and knowledge graphs;

[0012] Develop EPC knowledge sharing platform system.

[0013] Preferably, the graph database model layer design and data collection based on the knowledge type in EPC engineering management include:

[0014] Design the graph database model layer based on the knowledge types in EPC engineering management;

[0015] Collect large-scale multi-source data from relevant EPC projects, and perform pre-processing such as deletion, cleaning, and conversion of the collected raw data;

[0016] For non-text data, special processing of feature extraction is performed.

[0017] Preferably, the special processing of feature extraction for non-text data includes:

[0018] The attributes of drawings and presentation files are extracted and processed as their feature information; these processed non-text data will be linked to the corresponding entity nodes in the graph database as associated files, and the extracted feature text will be used as the attributes of the entity nodes. In this way, multi-source data will be uniformly converted into entity nodes represented by the extracted features.

[0019] Preferably, the training of specialized entity recognition and relationship extraction models based on entity and relationship information in EPC project management includes:

[0020] Based on the entity and relationship information in EPC engineering management, we train specialized entity recognition and relationship extraction models, and continuously iterate to improve the recognition accuracy of the models for EPC domain proper nouns, terminology, and key concepts.

[0021] Using the improved knowledge extraction model, knowledge is extracted from text information in the EPC engineering field to obtain a large number of knowledge triples and form a preliminary knowledge graph.

[0022] Preferably, the improved knowledge extraction model is used to extract knowledge from text information in the EPC engineering field, obtain a large number of knowledge triples, and form a preliminary knowledge graph. The specific processing flow of knowledge extraction and knowledge graph generation is as follows:

[0023] Connect to the EPC text database, write an automatic program, input text sentences into the knowledge extraction model line by line, and use the entity extraction model to first identify all entities in the sentence and output the entities and their types in the sentence. Then, input the text sentences and the entities output by the entity recognition model into the relationship extraction model, and output their relationship types. Finally, automatically store the obtained "head entity-relationship-tail entity" triple knowledge into the graph database to form a preliminary knowledge graph.

[0024] Preferably, the designing and training of the EPC knowledge fusion model and performing the entity alignment operation include:

[0025] The entities representing the same entity in the obtained preliminary EPC management knowledge graph are unified and merged to eliminate redundancy, ensure the consistency of entities, and provide support for subsequent knowledge reasoning.

[0026] Preferably, the designing and training of the knowledge reasoning model and performing the knowledge completion operation include:

[0027] Design and train knowledge reasoning models, perform knowledge completion operations, and automatically perform reasoning and prediction on the obtained fused knowledge, so as to infer potential EPC knowledge without increasing the burden of data collection and realize knowledge generation and transfer.

[0028] Preferably, the inferred data and the actual data are marked and distinguished:

[0029] The set of triples completed by knowledge is imported into the knowledge base, and an inferred attribute is set for each relationship obtained by reasoning completion, with a value of true, to mark these data as inferred data; through this attribute, the system can effectively distinguish inferred data from real data to avoid confusion between the two.

[0030] Preferably, the construction of a question-answering system framework based on a large model and knowledge graph includes:

[0031] Build a question-answering system framework based on a large model and knowledge graph, and configure a large model base that can query the EPC knowledge base and answer questions;

[0032] The output quality is optimized through Prompt engineering, and knowledge search templates and experience recommendation templates are designed to achieve accurate query, inference and expansion of knowledge.

[0033] Preferably, a large model base capable of querying the EPC knowledge base and answering questions is configured, including:

[0034] Select and determine the model architecture and version to be used, and configure relevant parameters;

[0035] The development of the EPC knowledge sharing platform system includes:

[0036] Set up source data storage module, knowledge extraction module, knowledge fusion module, knowledge reasoning module, knowledge graph visualization module, knowledge question and answer module, data security management module and system management module. Through the design and development of these modules, build a complete knowledge sharing platform for the EPC field.

[0037] The present invention has the following beneficial effects:

[0038] 1. This invention integrates multiple heterogeneous data sources and utilizes automated data collection and information extraction technologies to comprehensively collect professional knowledge, technical standards, regulations, policies, and industry best practices related to EPC (engineering, procurement, and construction) projects, ensuring the integrity and systematic nature of the knowledge system. Through structured processing and categorized management of various knowledge documents, it achieves comprehensive coverage of all types of information throughout the EPC project process, enabling timely and systematic compilation and aggregation of knowledge from project design and procurement to construction, operations, and maintenance. This provides project managers with more comprehensive and accurate knowledge support, improving decision-making efficiency and execution.

[0039] 2. The present invention uses natural language processing (NLP) and machine learning (ML) algorithms to automatically fuse and integrate knowledge from different professional fields. In EPC projects, there are large differences and independence between the knowledge of different professions. By constructing a multi-level semantic association map, the present invention can efficiently integrate various types of professional knowledge of structured and unstructured data across fields to ensure the accuracy and consistency of the data. The system uses high-precision semantic understanding and reasoning capabilities to eliminate barriers between professional knowledge, provide effective technical support for cross-professional collaboration, and improve the sharing and utilization efficiency of information.

[0040] 3. This invention combines machine learning with deep learning models to achieve automatic updating and self-growth of the knowledge base. The system can dynamically expand and modify the knowledge base in real time based on new project data, industry trends, technological developments, and regulatory updates. Through adaptive learning mechanisms, the system can identify outdated information and automatically eliminate or replace knowledge, while also incorporating the latest advances in emerging technologies, management methods, and relevant regulations to ensure the knowledge base remains efficient and advanced. This self-growth feature not only reduces the need for manual intervention but also significantly improves the system's responsiveness to unexpected issues in complex engineering projects.

[0041] 4. By integrating a large-scale language model (LLM), the present invention demonstrates an extremely high level of intelligence in knowledge query and interaction in EPC project management. Through deep semantic understanding and natural language generation technology, LLM can accurately understand the user's query intention and provide precise answers or recommendations. Whether facing the diverse needs of project managers, engineers or other professionals, the system can provide customized and personalized knowledge services. This model not only supports a variety of natural language queries, but also has cross-domain and multi-dimensional intelligent recommendation functions. It can adapt to complex and dynamic application scenarios in the EPC field, such as project progress management, risk assessment, resource allocation optimization, etc., significantly improving the efficiency of knowledge management and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The present invention will be further described below with reference to the accompanying drawings and examples.

[0043] Figure 1 This is a flowchart of a method for constructing a knowledge sharing platform for the EPC field provided by an embodiment of the present invention.

[0044] Figure 2 This is a flowchart of the professional knowledge conversion, processing and storage of EPC field data provided by an embodiment of the present invention.

[0045] Figure 3 Schematic diagram of the structure of the question-answering system based on the big model + knowledge graph provided in an embodiment of the present invention.

[0046] Figure 4 This is a system architecture diagram of the EPC knowledge sharing platform provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0048] See also Figure 1 The present invention provides a method for constructing a knowledge sharing platform for the EPC field, which can better share and utilize the internal knowledge assets in the EPC engineering field, reduce the human and material resource investment caused by knowledge management, and achieve cost reduction and efficiency improvement.

[0049] The method comprises:

[0050] S1, based on the knowledge types in EPC engineering management, designs the graph database model layer, collects multi-source data of relevant EPC engineering projects on a large scale, performs pre-processing such as deletion, cleaning and conversion on the collected raw data, and performs specialized processing such as feature extraction on non-text data.

[0051] In S2, specialized entity recognition and relationship extraction models are trained based on the entity and relationship information in EPC project management. Through continuous iteration, the models' recognition accuracy for EPC domain terms, terminology, and key concepts is improved. Using the improved knowledge extraction model, knowledge is extracted from the EPC engineering domain text information in S1, resulting in a large number of knowledge triples and forming a preliminary knowledge graph.

[0052] S3 designs and trains the EPC knowledge fusion model, performs entity alignment operations, and unifies and merges the entities representing the same entity in the preliminary EPC management knowledge graph obtained in S2 to eliminate redundancy, ensure entity consistency, and provide support for subsequent knowledge reasoning.

[0053] S4 designs and trains the knowledge reasoning model, performs knowledge completion operations, and automatically performs reasoning and prediction on the fused knowledge obtained in S3, so as to infer the potential EPC knowledge without increasing the burden of data collection and realize knowledge generation and transfer.

[0054] S5 builds a question-answering system framework based on a large model and knowledge graph, configures a large model base that can query the EPC knowledge base and answer questions, optimizes output quality through Prompt engineering, and designs knowledge search templates and experience recommendation templates to achieve accurate query, inference, and expansion of knowledge.

[0055] S6, develop the EPC knowledge sharing platform system, set up source data storage module, knowledge extraction module, knowledge fusion module, knowledge reasoning module, knowledge graph visualization module, knowledge question and answer module, data security management module and system management module. Through the design and development of these modules, a complete knowledge sharing platform for the EPC field can be built.

[0056] Example 2:

[0057] To facilitate understanding of this embodiment, a method for constructing a knowledge sharing platform for the EPC field provided by an embodiment of the present invention is first described in detail. Figure 1 As shown, the method includes the following steps:

[0058] S1, based on the knowledge types in EPC engineering management, designs the graph database model layer, collects multi-source data of relevant EPC engineering projects on a large scale, performs pre-processing such as deletion, cleaning and conversion on the collected raw data, and performs specialized processing such as feature extraction on non-text data.

[0059] Specifically, S1 includes the following steps:

[0060] S11, based on the knowledge type in EPC engineering management, determine the data category to be collected, and combine it with the actual usage scenario, with query decision-making as the core, to design the graph database model structure.

[0061] Specifically, as shown in Appendix 1, EPC management knowledge types cover a wide range, including but not limited to the following categories: project professional and technical data, involving core technical documents such as design drawings, technical specifications, and construction process flows; project organizational management data, including the organizational structure, role allocation, task schedule, and management processes of all parties involved in the project; construction management data, covering construction plans, site management, progress tracking, quality control, and other content; safety and environmental requirements, involving safety standards, environmental protection measures, and health management regulations that must be adhered to during project construction; and material procurement data, covering key aspects such as procurement plans, supplier information, procurement contracts, and logistics and transportation. Furthermore, based on the actual needs of the project, knowledge types can be further expanded to include other related areas such as financial management, risk control, and contract management to ensure the comprehensiveness and adaptability of the knowledge base.

[0062] Table 1 EPC management knowledge types

[0063]

[0064] Specifically, the graph database's schema design is guided by practical usage scenarios, focusing on optimizing query decisions. When designing entity classifications based on EPC domain knowledge types, the importance of each node category, relationship, and node attribute in the traversal and retrieval process is fully considered. Based on the principles of maximizing read and write efficiency, optimizing relationship queries, and minimizing matching depth, we strive to increase the number of node categories and reduce node attributes to improve database query performance. Furthermore, to ensure easy identification and structural consistency of database fields, entity naming uses unified professional terminology and maintains a consistent naming style.

[0065] The schema structure of a graph database can include entities such as "project management," "construction tasks," and "material procurement," while relationships can include "responsible," "association," and "procurement." These entities and relationships constitute the core structure of the graph database and provide target references for subsequent entity extraction.

[0066] S12, through API interface data collection tools or manual collection methods, collect relevant EPC engineering project data on a large scale. The data formats include but are not limited to text data, PDF text that can be recognized by OCR, images with explanatory information, videos, audio, professional engineering drawings, presentations, etc., to ensure the coverage and extensiveness of the data.

[0067] S13, pre-processing operations are performed on the collected original document data, specifically including steps such as data deletion, cleaning, and format conversion. Delete irrelevant or duplicate document entries, such as repeatedly uploaded construction plans, design change reports, or accompanying materials unrelated to the EPC project, such as supplier promotional materials or irrelevant contract attachments; perform data cleaning operations to correct format errors and typesetting problems in project management documents, or remove noise information, such as redundant headers and footers, marking lines in construction drawings, advertisements, or irrelevant watermarks, to improve the standardization and readability of the data. For example, correct the formatting problems of tables in the construction progress report to ensure that the data can be correctly parsed by the system in subsequent processing. On this basis, perform format conversion operations to uniformly convert document materials in different formats (such as construction plans in PDF format, equipment purchase lists in Word format, and construction logs in image format) into processable standardized text formats or structured data. For example, use OCR technology to extract text information from the completion report in PDF to facilitate the subsequent training and application of knowledge extraction models.

[0068] S14: For non-text data such as images, videos, and audio, information is extracted and organized through specialized processing tools or algorithms.

[0069] Specifically, for image data, image recognition technology is used to extract visual feature information related to graph database entities; for video data, key information is extracted through key frame extraction, speech recognition, and video content analysis; for audio data, speech recognition or audio feature extraction technology is used to convert the audio content into text; the attributes of files such as drawings and presentations are extracted and processed as their feature information. After processing, these non-text data will be linked to the corresponding entity nodes in the graph database as associated files, and the extracted feature text will be used as the attributes of the entity nodes. In this way, multi-source data are uniformly converted into entity nodes represented by extracted features, ensuring that these multi-source data can be effectively associated with related entities in the subsequent retrieval and analysis process, further enhancing the representation ability and query accuracy of the graph database. All pre-processed data can be uniformly stored in a local database or backed up to a cloud database.

[0070] In S2, specialized entity recognition and relationship extraction models are trained based on the entity and relationship information in EPC project management. Through continuous iteration, the models' recognition accuracy for EPC domain terms, terminology, and key concepts is improved. Using the improved knowledge extraction model, knowledge is extracted from the EPC engineering domain text information in S1, resulting in a large number of knowledge triples and forming a preliminary knowledge graph.

[0071] Specifically, S2 includes the following steps:

[0072] In S21, given the scarcity of specialized dictionaries in the EPC engineering field and the relative lack of annotated data within the field, we introduced the deep learning sequence annotation model BERT+BiLSTM+CRF during entity extraction. Based on the predefined graph database entities from S11, we prepared annotated EPC text training data and trained the BERT+BiLSTM+CRF model.

[0073] Specifically, the BERT (Bidirectional Encoder Representations from Transformers) module is used to convert annotated text into a vector sequence that can be recognized by a neural network. The vector not only contains the semantic information of each word, but also captures the contextual information of the entire sentence through BERT's bidirectional encoder, ultimately converting the original corpus into word vectors.

[0074] Among them, the vector sequence can be expressed as:

[0075]

[0076] Where, X i Encode the vector sequence corresponding to the i-th word in the sentence, represents the embedding vector of the word itself, Represents the position embedding vector of the word, Represents the sentence embedding vector of the word.

[0077] Specifically, the Bi-directional long short-term memory (BiLSTM) further processes the representation vector of each word output by BERT. Through the bidirectional LSTM structure, it captures more detailed contextual dependencies and extracts the linguistic characteristics of the text in context, thereby providing richer input features for the CRF layer.

[0078] Among them, the bidirectional LSTM structure can be expressed as:

[0079]

[0080] Where t represents the current input time step, xt Represents the t-th word vector in the input sequence, and the function LSTM() represents the calculation process using the LSTM unit. represents the hidden state of the reverse LSTM at the current time step, represents the hidden state of the reverse LSTM at the current time step, h t It is the fusion of BiLSTM's context representation of each word in the entire sentence.

[0081] Specifically, CRF (Conditional Random Field) uses conditional random fields to learn the relationships between labels, maximize the conditional probability of the entire label sequence, and provide a globally optimal label sequence. This improves the accuracy of the result predictions output by BiLSTM and enhances the accuracy of the model in identifying specific entities in the EPC management field.

[0082] Among them, the conditional probability can be expressed as:

[0083]

[0084] Where P(Y|X) is the conditional probability of predicting the label sequence Y given the input sequence X, score(X,Y) is the score of the given label sequence Y, ∑ Y′∈y exp(score(X,Y′)) is the sum of the scores of all possible label sequences for normalization.

[0085] S22, considering the characteristics of EPC management field texts with multiple entity types and complex relationships, the improved BERT-BiGRU-Attention model is used to automatically extract the relationships between identified entities.

[0086] Specifically, the BERT module is used to convert annotated text into a vector sequence that can be recognized by the neural network. The vector not only contains the semantic information of each word, but also captures the contextual information of the entire sentence through BERT's bidirectional encoder, ultimately converting the original corpus into word vectors.

[0087] Specifically, the Bidirectional Gated Recurrent Unit (BiGRU) module is used as a sequence modeling tool to capture contextual features in text. The BiGRU module is similar in principle to the BiLSTM, but has a simpler structure.

[0088] Specifically, the Attention module is used to improve the model's ability to focus on important features. For EPC relationship extraction tasks, certain words or phrases are more important than others in determining relationships between entities. The Attention mechanism automatically adjusts the model's attention distribution by calculating the correlation between different time steps (or words), helping the model focus on the words most relevant to relationship determination.

[0089] Among them, the Attention mechanism layer can be expressed as:

[0090] e t =f(Query,Key t );

[0091]

[0092] c=∑ t α t h t ;

[0093] Where, e t is the attention score, f(Query,Key t ) is a similarity measurement function used to measure the matching degree between the query and the key, where Query is the query vector, Key t is the hidden state h at the current time step t in the BiGRU module t , α t represents the normalized attention score, that is, the word vector attention weight at the current time step t, and c represents the entire sentence vector output after the Attention module.

[0094] S23, after training the entity extraction model and relationship extraction model of S21 and S22, obtains the EPC domain knowledge extraction model, and performs knowledge extraction on the large amount of text data converted into the format obtained in S13.

[0095] The knowledge extraction includes: automatic data input, running the BERT+BiLSTM+CRF entity extraction model, running the BERT-BiGRU-Attention relationship extraction model, and storing the output triplet knowledge.

[0096] The specific processing flow of knowledge extraction and knowledge graph generation is as follows: connect to the EPC text database, write an automatic program, input the text sentences into the knowledge extraction model line by line, first use the entity extraction model to identify all entities in the sentence, output the entities in the sentence and their types, then input the text sentences and the entities output by the entity recognition model into the relationship extraction model, and output their relationship types, finally automatically store the obtained "head entity-relationship-tail entity" triple knowledge into the graph database to form a preliminary knowledge graph.

[0097] S3 designs and trains the EPC knowledge fusion model, performs entity alignment operations, and unifies and merges the entities representing the same entity in the preliminary EPC management knowledge graph obtained in S2 to eliminate redundancy, ensure entity consistency, and provide support for subsequent knowledge reasoning.

[0098] Specifically, S3 includes the following steps:

[0099] S31, before building the entity alignment model, in order to enable the model to more effectively handle different types of attributes and relationships and address the processing requirements of different attribute types (such as text, numerical values, etc.), this method is designed to perform entity alignment preprocessing operations. According to the different attribute types, the complex information of the knowledge graph is decomposed into more targeted substructures, and the knowledge graph knowledge formed in S23 is divided into: triples containing the "name" attribute, triples containing attribute values ​​as descriptive text, triples of numerical types, and triples containing relationship triples. Then, according to the characteristics of different knowledge graph subgraphs, different node feature extraction and vectorization methods are adopted.

[0100] Specifically, for feature extraction of name attribute subgraphs and descriptive text attribute subgraphs, text embedding models (such as BERT and Word2Vec) are used to convert entity names into vector representations; for feature extraction of numerical attribute subgraphs, standardization or numerical embedding of numerical attributes is adopted to convert numerical attributes into vector representations; for feature extraction of relationship subgraphs, graph structure feature extraction methods are used, such as GNN to generate node embedding vectors.

[0101] S32, build the AttrGNN entity alignment model and train it to form a knowledge fusion model.

[0102] In particular, the present invention aims to integrate information from one or more knowledge sources or knowledge bases into a unified and complete knowledge system through knowledge fusion, thereby removing redundant information from the knowledge graph. Specific operations of knowledge fusion include entity alignment, relationship merging, and attribute fusion. However, since the knowledge extraction model defined by the present invention may cause entity misalignment during the extraction process, a special entity alignment model needs to be constructed to ensure the consistency of knowledge. For relationships and attributes, the processing effect of the knowledge extraction model defined by the present invention is relatively ideal, so there is no need for additional relationship merging and attribute fusion operations.

[0103] In particular, the knowledge fusion in the present invention mainly refers to entity alignment, which can also be called entity merging.

[0104] Specifically, AttrGNN (Attribute Graph Neural Network) uses the GNN aggregation mechanism to simultaneously consider graph structure information and node attribute information to more comprehensively represent nodes. It uses an attention mechanism to assign different weights to different attributes of entities, effectively alleviating severe dataset bias and achieving significant improvements and better performance. The AttrGNN model's modeling process includes: training sample data preprocessing (subgraph partitioning, node feature extraction and vectorization), attribute weighted aggregation, adjacency feature information aggregation, multi-layer graph convolution, feature combination and similarity calculation, training optimization based on a loss function, and alignment decision-making.

[0105] Among them, the core mechanism of AttrGNN is expressed as:

[0106]

[0107] in, is the feature representation of entity e in the first layer, α ej is the attention weight, which represents the attention score between node j and its neighbors, Msg(j) is the attribute and graph structure information collected from node j, and σ is a nonlinear activation function (such as ReLU or Sigmoid); is the feature representation of entity e in the second layer, W2 is a learnable weight matrix used to linearly transform the aggregated information, and mean represents the average aggregation of the features of node e itself and the features of all its neighbors N(e).

[0108] The loss function for model training uses the following margin-based ranking loss formula:

[0109]

[0110] Where P represents the positive sample set, including known matching entity pairs (e1, e2), N represents the negative sample set, including unmatched entity pairs (e1′, e2′), is the distance between the embedding vectors of two entities e1 and e2, using Euclidean distance or cosine similarity. γ is a marginal hyperparameter that controls the minimum distance difference between positive and negative samples, [·] + Represents a non-negative function, which means that if the value in the brackets is less than zero, it is taken to be zero; otherwise, it retains its original value.

[0111] S33, based on the entity alignment model trained in S31, makes entity alignment decisions for all knowledge triples in S23. The specific process includes: triple data preprocessing (subgraph partitioning, node feature extraction and vectorization), using the AttrGNN model to obtain entity alignment decision output, screening alignment targets according to the set decision threshold, and updating the knowledge graph.

[0112] Specifically, the node feature extraction and vectorization steps must be consistent with the methods used in S31 model training.

[0113] Specifically, the decision threshold is the threshold set for the similarity of each pair of nodes output by the model (such as 0.9). When the similarity of two nodes exceeds this threshold, they will be regarded as the same entity and merged.

[0114] Specifically, updating the knowledge graph requires using database operation languages ​​such as Cypher to process the preliminary knowledge graph obtained in S23, delete duplicate nodes, and retain the remaining relevant information.

[0115] S4 designs and trains the knowledge reasoning model, performs knowledge completion operations, and automatically performs reasoning and prediction on the fused knowledge obtained in S3, so as to infer the potential EPC knowledge without increasing the burden of data collection and realize knowledge generation and transfer.

[0116] Specifically, S4 includes the following steps:

[0117] S41. In the knowledge graph, it is inevitable that some knowledge will be missing (such as some relationships are not clearly recorded). In order to promote the cross-domain application and transfer of knowledge and promote knowledge innovation, this method designs a model that predicts missing entity relationships and can perform knowledge reasoning and completion. This method is based on TransH, optimizes model training parameters, and constructs and trains reasoning and completion models. Specifically, the steps of model construction and training include: constructing a training data set (a large amount of head entity, relationship, tail entity triple training data, entity and ID mapping data, and building relationship and ID mapping data); setting loading training parameters (including data set path, batch number, positive and negative sample sampling mode, whether to filter invalid samples, the number of negative samples in each batch, and the number of negative relationships in each batch); defining the TransH model structure (configuring parameters such as embedding dimension, norm, and whether to regularize the vector); defining the loss function; iteratively updating the training model; and model testing.

[0118] Specifically, the TransH model defines a hyperplane determined by the normal vector w for each relation r in the triple (h, r, t). The head entity h and the tail entity t are projected onto the hyperplane of the relation r:

[0119]

[0120] Among them, h ⊥ , t ⊥ are the projections of entities h and t on the hyperplane of relation r, respectively.

[0121] The core goal of TransH is to make the distance between the projected vectors in the triple (h, r, t) as close as possible:

[0122] d(h,r,t)=||h ⊥ +rt ⊥ ||;

[0123] Specifically, the marginal loss function (Margin-based Ranking Loss) is used to optimize the model during model construction:

[0124] L=∑ 正样本 ∑ 负样本 max(0,γ+d(h,r,t)-d(h′,r,t′));

[0125] S42: Perform link prediction and generate completion results. The prepared prediction data (incomplete triples consisting of head entities and relationships, but missing tail entities) is fed into the trained model. The model predicts multiple candidate tail entities for each triple and assigns a score to the generated complete triples. Based on a set score threshold, combined with review by domain experts or real-world data verification, reasonable results are screened, ultimately generating a complete set of triples after knowledge completion.

[0126] S43 imports the completed knowledge triples into the knowledge base. At the same time, an inferred attribute is set for each completed relationship, with a value of true, marking this data as inferred. This attribute allows the system to effectively distinguish inferred data from real data, preventing confusion between the two. Future queries can flexibly filter or focus on inferred data based on the inferred attribute, allowing users to view only completed or migrated knowledge, ensuring transparent and reliable data management.

[0127] S5 builds a question-answering system framework based on a large model and knowledge graph, configures a large model base that can query the EPC knowledge base and answer questions, optimizes output quality through Prompt engineering, and designs knowledge search templates and experience recommendation templates to achieve accurate query, inference, and expansion of knowledge.

[0128] Specifically, S5 includes the following steps:

[0129] S51, using LangChain, builds a question-answering system framework based on large models + knowledge graphs to achieve intelligent question-answering of professional knowledge.

[0130] Specifically, LangChain is a development framework for large language models that supports large language models to access external data sources such as knowledge bases, allowing them to receive natural language questions input by users and output results based on knowledge base queries.

[0131] S52: Build a general large-scale language model (LLM) foundation to parse user-entered questions related to EPC knowledge. This foundation should have the following functions: extract key entities and relationships in user questions; query and search the knowledge base for relevant nodes and relationships; organize and summarize the results of knowledge base queries; and generate concise, easy-to-understand natural language answers. Specific steps include: selecting and determining the model architecture and version to use (such as Tongyi Qianwen-Turbo, Tongyi Qianwen-Max, ChatGPT-3.5, ChatGPT-4, etc.) and configuring relevant parameters (such as whether to enable historical sessions, maximum context length, response speed optimization, etc.).

[0132] S53, to improve the output quality of the knowledge question-answering system, high-quality prompt word templates are designed through the Prompt project to optimize knowledge query results. The Prompt project design includes the following: presetting background information and query context for knowledge base searches to control output information density and answer length; guiding the model to focus on knowledge queries and knowledge-experience analogies to ensure logical and hierarchical answers; and setting prompt templates for different contextual modes to improve the practicality and accuracy of answers.

[0133] Specifically, a knowledge search prompt template is created based on the actual usage scenarios of EPC knowledge. This template is used in knowledge query scenarios, prompting the large model to explore relevant relationships as much as possible when querying the knowledge graph, ensuring the comprehensiveness of the query information. By optimizing output through natural language, the model not only returns direct answers but also provides logically coherent contextual explanations, allowing users to obtain more in-depth and relevant answers when dealing with complex problems.

[0134] Specifically, based on the actual use cases of EPC knowledge, an experience recommendation prompt template is set. This template is used to identify user intent and provide experience-based recommended answers for questions focusing on analogy, transfer, or summary. When the model lacks a direct answer, it automatically searches the knowledge base for similar knowledge and generates experience-based answers. This setting is suitable for query-related knowledge that does not exist in the knowledge base, enabling knowledge expansion and recommendation, and providing valuable insights.

[0135] S6, develop the EPC knowledge sharing platform system, set up source data storage module, knowledge extraction module, knowledge fusion module, knowledge reasoning module, knowledge graph visualization module, knowledge question and answer module, data security management module and system management module. Through the design and development of these modules, a complete knowledge sharing platform for the EPC field can be built.

[0136] Specifically, S6 includes the following steps:

[0137] S61, design and development of source data storage module. This module needs to have the following functions: establish an EPC field source data upload interface to support the upload and centralized management of source data and its ancillary information; build structured and unstructured databases based on different types of source data to ensure efficient reading and writing of data in various formats; deploy automated data cleaning and conversion models to achieve key operations such as data deduplication, format unification, and character encoding standardization.

[0138] S62, design and development of knowledge extraction module. This module must have the following functions: provide management functions for knowledge extraction models, support model training, optimization and updating; implement real-time monitoring functions for the extraction process, display automatically extracted entities, relationships and their confidence scores, and provide manual correction or deletion functions for extracted triples to ensure the accuracy of knowledge extraction results.

[0139] S63, design and development of knowledge fusion module, which must have the following functions: management and update functions of the knowledge fusion model, supporting continuous training and optimization of the model; visual presentation of the fusion process, including a progress monitoring interface that displays the progress of each entity merger and the associated information of the merged entity nodes; redundancy elimination and consistency verification, after the fusion is completed, based on the set similarity threshold, automatically detect redundant entities and inconsistent relationships in the graph, and prompt the user to make manual corrections; provide a manual intervention interface to support users to mark entity relationships or alignment rules that need to be retained to meet personalized adjustment needs.

[0140] S64, design and development of knowledge reasoning module, this module needs to have the following functions: provide management and update functions of knowledge reasoning model, support continuous training and optimization of the model; be able to visualize the knowledge graph structure before and after knowledge reasoning; the node information of the knowledge graph after knowledge reasoning has the attribute of inferred, marking the knowledge relationship derived by the reasoning model; provide a manual intervention interface to support users to actively delete the nodes derived by reasoning and other operations.

[0141] S65, design and development of knowledge graph visualization module, which needs to have the following functions: knowledge graph display function, using visualization tools (such as @antv / g6) to realize the visualization of graph nodes and relationships, and support interactive operations such as node dragging, zooming, etc.; query and filtering function, designing query interface, supporting query and filtering based on node and relationship attributes, such as "displaying all relationships and nodes related to a specific EPC node"; path search function, integrating path search algorithm, users can enter two nodes, and the system will display the relationship path between them for knowledge link analysis; hot relationship and association display function, providing graph heat map, showing high-frequency association relationships, so that users can identify key nodes in knowledge sets.

[0142] S66, design and development of knowledge question and answer module, which needs to have the following functions: user and large model dialogue function, with simple dialog box design and smooth response, and a slightly delayed transmission method to enhance the real dialogue experience; session storage function, generating a unique session ID for each question and answer, and storing the user's questions and the model's answers in the database in chronological order to facilitate subsequent query and display; popular question and answer module, displaying common or highly visited questions; feedback mechanism, allowing users to evaluate answers to support subsequent optimization and improvement of the large model prompt words.

[0143] S67, design and development of data security management module, which must have the following functions: permission management function, building a hierarchical permission system based on user roles (such as administrators, ordinary users), ensuring refined management of system login and access control to limit access rights to sensitive information; data operation log and audit function, comprehensively recording user operation logs, including data access, model use, question and answer records, etc., to ensure the traceability of operations and meet audit and compliance requirements.

[0144] S68, system management module design and development, this module must have the following functions: user information management function, create a detailed personal profile for each user, including user name, contact information, role permissions and other information; provide a personal log query interface, support users to filter logs according to time, operation type and other conditions, convenient for users to query and review their operation records; record all system-level events (such as system startup, module update, fault report, etc.) to ensure that administrators can monitor the system operation status in real time and comprehensively, and improve the visualization and transparency of system management.

[0145] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0146] The example steps of S23 knowledge extraction and knowledge graph generation are as follows:

[0147] The following is a brief introduction to the EPC project of the micro-jacking sewage interception project of the Gezhouba Wharf. The knowledge extraction model defined in this invention is used to extract knowledge from this project overview: The Gezhouba Wharf sewage interception project is located in Xiling District, Yichang City. The starting point is at the intersection of Zhenping Road and Shangdaodi Road. The project is located along Shangdaodi Road to Gezhouba Park. This project is located at KW0+126-WK0+314 of the Shangdaodi Road section of Gezhouba Wharf in Yichang City. In order to avoid damage caused by pipeline excavation construction, This section of the sewage interception pipeline, between inspection wells W5-1 and W10, will be constructed using micro-pipe jacking technology. The pipeline will be buried deep underground, with a designed bottom elevation of 65.99 to 66.27 meters and a buried depth of 4.07 to 7.47 meters. The distance between the two inspection wells is designed to be between 40 meters. The diameter of the pipes W5-1 to W10 is DN500, with a total length of 188 meters. A total of six micro-pipe jacking working wells (receiving wells) will be located between W5-1 and W10. Surrounding the area are hazardous factors such as walls, auto repair shops, green trees, and underground pipelines. The construction period is scheduled from July 5, 2021, to October 18, 2021.

[0148] Step 1: Automatically input text sentence by sentence into the entity recognition model;

[0149] Step 2: The entity recognition model outputs the entities and attributes contained in the sentence. Geographic location (Xiling District, Yichang City, Zhenping Road, Shangdaodi Road, Gezhouba Park); Project name (Gezhouba Wharf Sewage Interception Project, Sewage Interception Pipeline); Date (July 5, 2021, October 18, 2021); Technology (Micro-Pipe Jacking Technology); Elevation (65.99m, 66.27m); Buried depth (4.07m, 7.47m); Pipeline length (188m); Inspection wells (W5-1, W10);

[0150] Step 3: Relationship extraction and triple output. (Gezhouba Wharf Sewage Interception Project, located in Xiling District, Yichang City); (Gezhouba Wharf Sewage Interception Project, using micro-pipe jacking technology); (Gezhouba Wharf Sewage Interception Project, starting point, intersection of Zhenping Road and Shangdaodi Road); Gezhouba Wharf Sewage Interception Project, end point, Gezhouba Park); (Gezhouba Wharf Sewage Interception Project, pipeline length, 188 meters); (Gezhouba Wharf Sewage Interception Project, construction start date, July 5, 2021); (Gezhouba Wharf Sewage Interception Project, construction end date, October 18, 2021); (W5-1, to, W10); (W5-1, burial depth, 4.07 meters);

[0151] Step 4: Store the triples in a graph database. The generated triples are stored in a graph database such as Neo4j or NabulaGraph. In a graph database, the head and tail entities are nodes, and the relationships are edges, thus forming a preliminary knowledge graph.

[0152] The example steps for entity alignment of the S33 knowledge triples are as follows:

[0153] The following are several triplets obtained according to the knowledge extraction model. These triplets are fused using the entity alignment model defined in the present invention: (W5-1 working well, location, Dongshan Avenue), (W5-1 working well, location, Dongshan Avenue area), (W-8 working well, pipe diameter, 500 mm), (W-8 well, pipe diameter, DN500), (Yichang sewage treatment plant renovation project, construction technology, excavation construction), (Yichang sewage treatment project, construction technology, excavation construction technology).

[0154] Step 1: Divide the knowledge graph into subgraphs according to the different attribute types of the entity. Divide the above triples into:

[0155] Location submap: (W5-1 working well, location, Dongshan Avenue), (W5-1 working well, location, Dongshan Avenue area);

[0156] Numerical attribute subgraph: (W-8 working well, pipe diameter, 500 mm), (W-8 working well, pipe diameter, DN500);

[0157] Construction technology attribute sub-graphs: (Yichang Sewage Treatment Plant Renovation Project, Construction Technology, Excavation Construction), (Yichang Sewage Treatment Plant Renovation Project, Construction Technology, Excavation Construction Technology);

[0158] Step 2: Node feature extraction, using different methods to extract and vectorize features for entities in each subgraph.

[0159] Dongshan Avenue: [0.21, 0.34, 0.11, ..., 0.43];

[0160] Dongshan Avenue area: [0.22, 0.32, 0.10, ..., 0.42];

[0161] 500mm: [0.5, 0.22, 0.61];

[0162] DN500: [0.5, 0.51, 0.48];

[0163] Excavation method: [0.18, 0.24, 0.33, ..., 0.50];

[0164] Excavation construction technology: [0.17, 0.25, 0.32, ..., 0.49];

[0165] Step 3: Generate node feature matrix and adjacency matrix. Node feature matrix: [[0.21, 0.34, 0.11, ..., 0.43], [0.22, 0.32, 0.10, ..., 0.42], [0.18, 0.24, 0.33, ..., 0.50], [0.17, 0.25, 0.32, ..., 0.49], [0.5, 0.22, 0.61], [0.5, 0.51, 0.48]], adjacency matrix: [[0, 1, 0, 0, 0, 0], [1, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0], [0, 0, 1, 0, 0], [0, 0, 0, 0, 0, 1], [0, 0, 0, 0, 1, 0]];

[0166] Step 4: Input the node feature matrix and adjacency matrix into the trained AttrGNN entity alignment model and output the entity pair similarity.

[0167] Similarity between Dongshan Avenue and Dongshan Avenue area: 0.95;

[0168] Similarity between 500mm and DN500: 0.94;

[0169] Similarity between excavation construction and excavation construction technology: 0.93;

[0170] Step 5: Decision-making for entity alignment. Based on the set similarity threshold of 0.9, the similarity results were evaluated: Dongshan Avenue and Dongshan Avenue Area exceeded the threshold and were determined to be the same entity; 500 mm and DN500 had similarities above the threshold and were determined to be the same entity; and excavation construction and excavation construction technology had similarities above the threshold and were determined to be the same entity.

[0171] Step 6: Update the knowledge graph and merge instance triples into entities.

[0172] The example steps of the reasoning completion steps S41 and S42 are as follows:

[0173] Currently, there are existing knowledge triplets regarding civil engineering foundation pit excavation, as well as triplets regarding key monitoring items for regulators, such as (civil engineering foundation pit, support method, steel support), (civil engineering foundation pit, construction method, layered excavation), (civil engineering foundation pit, key monitoring points, groundwater level), (civil engineering foundation pit, key monitoring points, building settlement), and (civil engineering foundation pit, risk points, groundwater leakage). Based on this existing knowledge, we hope to use the knowledge completion model obtained from S41 to perform knowledge transfer reasoning and completion, generating relevant knowledge about water pipeline construction foundation pits.

[0174] The first step is to construct a dataset based on the reasoning completion goal. This dataset consists of multiple incomplete triples consisting of only the head entity and related relationships of the water pipeline construction pit, but missing the tail entity. For example, (pipeline pit, support method, ...), (pipeline pit, monitoring points, ...), (pipeline pit, risk points, ...).

[0175] In the second step, the incomplete triplet dataset is input into the model obtained in S41, and the triplet prediction results are output and the scores are calculated. The output format is as follows:

[0176] (Pipeline foundation pit, support method, steel support) score: 0.30;

[0177] (Pipeline foundation pit, support method, anchor support) score: 0.83;

[0178] (Pipeline foundation pit, monitoring points, building settlement) score: 0.40;

[0179] (Pipeline foundation pit, monitoring points, groundwater level) score: 0.26;

[0180] (Pipeline foundation pit, key monitoring points, air pollution) score: 0.86;

[0181] (Pipeline foundation pit, risk point, groundwater leakage) score: 0.25;

[0182] The third step involves screening and verification using thresholds and domain experts. Inference triples with scores below 0.3 are considered highly credible, while those with scores between 0.3 and 0.5 are considered relatively credible and require subsequent expert review. After screening and expert review, the following highly credible inference completion triples are retained as knowledge obtained through inference completion: (pipeline foundation pit, support method, steel support), (pipeline foundation pit, key monitoring points, groundwater level), and (pipeline foundation pit, risk point, groundwater leakage).

[0183] The fourth step is to store the filtered triples in the knowledge base and set the inferred attribute to true for each relationship obtained through inference to distinguish inferred data from real data.

Claims

1. A method for constructing a knowledge sharing platform for the EPC field, characterized by: The method comprises: Based on the knowledge types in EPC engineering management, carry out graph database model layer design and data collection; Based on the entity and relationship information in EPC engineering management, train specialized entity recognition and relationship extraction models; Design and train the EPC knowledge fusion model to perform entity alignment operations; Design and train knowledge reasoning models to perform knowledge completion operations; Build a question-answering system framework based on large models and knowledge graphs; Develop EPC knowledge sharing platform system.

2. The method for constructing a knowledge sharing platform for the EPC field according to claim 1, characterized in that: The graph database model layer design and data collection based on the knowledge type in EPC engineering management include: Design the graph database model layer based on the knowledge types in EPC engineering management; Collect large-scale multi-source data from relevant EPC projects, and perform pre-processing such as deletion, cleaning, and conversion of the collected raw data; For non-text data, special processing of feature extraction is performed.

3. The method for constructing a knowledge sharing platform for the EPC field according to claim 2, characterized in that: The special processing of feature extraction for non-text data includes: The attributes of drawings and presentation files are extracted and processed as their feature information; these processed non-text data will be linked to the corresponding entity nodes in the graph database as associated files, and the extracted feature text will be used as the attributes of the entity nodes. In this way, multi-source data will be uniformly converted into entity nodes represented by the extracted features.

4. The method for constructing a knowledge sharing platform for the EPC field according to claim 3 is characterized in that: The aforementioned training of specialized entity recognition and relationship extraction models based on entity and relationship information in EPC engineering management includes: Based on the entity and relationship information in EPC engineering management, we train specialized entity recognition and relationship extraction models, and continuously iterate to improve the recognition accuracy of the models for EPC domain proper nouns, terminology, and key concepts. Using the improved knowledge extraction model, knowledge is extracted from text information in the EPC engineering field to obtain a large number of knowledge triples and form a preliminary knowledge graph.

5. The method for constructing a knowledge sharing platform for the EPC field according to claim 4, characterized in that: The improved knowledge extraction model is used to extract knowledge from text information in the EPC engineering field, obtain a large number of knowledge triples, and form a preliminary knowledge graph. The specific processing flow of knowledge extraction and knowledge graph generation is as follows: Connect to the EPC text database and write an automatic program to input text sentences into the knowledge extraction model line by line. The entity extraction model first identifies all entities in the sentence and outputs the entities and their types in the sentence. Then, the text sentence and the entities output by the entity recognition model are input into the relationship extraction model, and their relationship types are output. Finally, the obtained "head entity-relationship-tail entity" triple knowledge is automatically stored in the graph database to form a preliminary knowledge graph.

6. The method for constructing a knowledge sharing platform for the EPC field according to claim 5, characterized in that: The design and training of the EPC knowledge fusion model and the execution of entity alignment operations include: The entities representing the same entity in the obtained preliminary EPC management knowledge graph are unified and merged to eliminate redundancy, ensure the consistency of entities, and provide support for subsequent knowledge reasoning.

7. The method for constructing a knowledge sharing platform for the EPC field according to claim 6, characterized in that: The design and training of the knowledge reasoning model and the execution of the knowledge completion operation include: Design and train knowledge reasoning models, perform knowledge completion operations, and automatically perform reasoning and prediction on the obtained fused knowledge, so as to infer potential EPC knowledge without increasing the burden of data collection and realize knowledge generation and transfer.

8. The method for constructing a knowledge sharing platform for the EPC field according to claim 7, characterized in that: Mark and distinguish the inferred data from the actual data: The set of triples completed by knowledge is imported into the knowledge base, and an inferred attribute is set for each relationship obtained by reasoning completion, with a value of true, to mark these data as inferred data; through this attribute, the system can effectively distinguish inferred data from real data to avoid confusion between the two.

9. The method for constructing a knowledge sharing platform for the EPC field according to claim 8, characterized in that: The question-answering system framework based on the large model and knowledge graph includes: Build a question-answering system framework based on a large model and knowledge graph, and configure a large model base that can query the EPC knowledge base and answer questions; The output quality is optimized through Prompt engineering, and knowledge search templates and experience recommendation templates are designed to achieve accurate query, inference and expansion of knowledge.

10. The method for constructing a knowledge sharing platform for the EPC field according to claim 8, characterized in that: Configure a large model base capable of querying the EPC knowledge base and answering questions, including: Select and determine the model architecture and version to be used, and configure relevant parameters; The development of the EPC knowledge sharing platform system includes: Set up source data storage module, knowledge extraction module, knowledge fusion module, knowledge reasoning module, knowledge graph visualization module, knowledge question and answer module, data security management module and system management module. Through the design and development of these modules, build a complete knowledge sharing platform for the EPC field.