A petrochemical industry dangerous chemical warehouse area knowledge graph construction method based on documents

By constructing a knowledge graph of hazardous chemical storage areas in the petrochemical industry and combining it with deep learning technology, the shortcomings of manual analysis in the handling of hazardous chemical accidents have been addressed, enabling more scientific knowledge storage and emergency response, and improving the efficiency and accuracy of accident handling.

CN117112805BActive Publication Date: 2025-11-07FUJIAN SPECIAL EQUIP TESTING RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311111690.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-11-07
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

Current technologies for handling hazardous chemical accidents rely on manual analysis, which is labor-intensive, prone to human error, and lacks high-quality public datasets, hindering the application of deep learning in this field.

Method used

A document-based knowledge graph construction method for hazardous chemical storage areas in the petrochemical industry is adopted. By constructing a schema layer and a data layer of the knowledge graph and combining deep learning technology, entity information is extracted from hazardous chemical standards and accident emergency documents to establish a comprehensive knowledge graph for the prevention and emergency response of hazardous chemical accidents.

Benefits of technology

It enables scientific analysis and emergency response to hazardous chemical accidents, improves accident handling capabilities and the visualization and decision-making efficiency of emergency resources, and promotes the intelligent and precise management of hazardous chemicals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117112805B_ABST
    Figure CN117112805B_ABST
Patent Text Reader

Abstract

The application provides a petrochemical industry dangerous chemical warehouse area knowledge graph construction method based on documents, wherein the method constructs a mode layer and a data layer of the knowledge graph based on dangerous chemicals and dangerous chemical accident emergency disposal files; the mode layer is derived from dangerous chemical standard digitization and emergency response business logic, first finds out comprehensive abstract concepts in the field of dangerous chemicals, then continuously refines upper layer concepts into more specific lower layer concepts from top to bottom, defines concept entities, attributes and hierarchical semantic relationships from top to bottom, and constructs an accurate and structured hierarchical concept system architecture; the data layer is derived from bottom to top, extracts entity information and semantic associations based on deep learning from dangerous chemical specification files, academic literature and accident cases, and uses a twin neural network-based entity alignment model to fuse multiple knowledge graphs to form a comprehensive dangerous chemical warehouse area knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of petrochemical industry dangerous chemicals, and particularly relates to a petrochemical industry dangerous chemicals warehouse area knowledge graph construction method based on documents. BACKGROUND

[0002] Any article with corrosive, natural, flammable, toxic, explosive and other properties, which is easy to cause personal injury and property damage during transportation, loading and unloading, and storage and preservation, and must be specially protected, is a dangerous goods. Dangerous goods have special physical and chemical properties. If not properly protected during transportation, accidents are likely to occur, and the consequences of the accidents are more serious than general vehicle accidents. Existing dangerous chemical accidents are handled manually, so the analysis and judgment of on-site risk hidden dangers are completed manually, which is a huge workload and is prone to human errors.

[0003] The occurrence of dangerous chemical accidents involves many factors, including the stability, physicochemical properties, toxicological properties, environmental hazard properties of dangerous chemicals, and human health effects, ecological effects, environmental effects, and social effects after the accidents. The data involved have problems such as wide information range, large data volume, complex data types, multiple data sources, and uneven quality. At the same time, there is a lack of complete knowledge information in the field of dangerous chemicals, and there is currently no high-quality public data set, which hinders scientific research progress and makes it difficult to apply new technologies such as deep learning in the field of dangerous chemicals.

[0004] With the help of knowledge graph and Chinese information processing technology, a reasoning and analysis model of dangerous chemical accidents can be constructed, which can not only establish a plan for the prevention of dangerous chemical accidents in advance, but also assist in establishing an emergency response plan to improve the handling capacity of dangerous chemical accidents.

[0005] A knowledge graph is a semantic network that reveals the relationship between entities. The nodes can represent related entities, and the edges represent the relationship between the two entities. Entities are connected to each other by existing relationships, forming a network that contains rich semantics. In a computer, the entities and their relationships in the real world are expressed, which helps to discover deep connections between entities and facilitates related reasoning and analysis. SUMMARY

[0006] To overcome the above problems, the present application aims to provide a petrochemical industry dangerous chemicals warehouse area knowledge graph construction method based on documents, which directly displays the results in the form of a knowledge graph, providing a new analysis tool and approach for accident cause analysis.

[0007] The application adopts the following scheme: a petrochemical industry dangerous chemical warehouse area knowledge graph construction method based on documents, characterized in that: the method constructs a mode layer and a data layer of the knowledge graph based on dangerous chemicals and dangerous chemical accident emergency disposal files;

[0008] The mode layer is based on the digitalization of dangerous chemical standards and emergency response business logic, first finds out the comprehensive abstract concepts in the field of dangerous chemicals, and then continuously refines the upper layer concepts into more specific lower layer concepts from top to bottom, defines the concept entities, attributes and hierarchical semantic relationships from top to bottom, and constructs an accurate and structured hierarchical concept architecture; through continuous induction, clustering and generalization processing of lower layer concepts and terms in the field knowledge, the abstract concept model of the upper level is comprehensively obtained;

[0009] The data layer is from bottom to top, and different data of dangerous chemical specification files, academic literature and accident cases are aligned and fused according to the characteristics of dangerous chemicals and the measures adopted in the accident response based on deep learning to extract entity information and semantic association, establish the mapping between specific elements and concept nodes, form the mapping from the mode layer to the data layer, and use the entity alignment model based on the twin neural network to fuse multiple knowledge graphs to form a comprehensive dangerous chemical warehouse area knowledge graph.

[0010] Further, academic literature analysis and research are carried out to find the specification files of dangerous chemicals, including the following steps:

[0011] 2.1 Determine the scope and target of the study: read the academic literature to clearly understand the content and scope of the specification files of dangerous chemicals;

[0012] 2.2 Collect specification files: use academic search engines, library resources, government agency websites and other channels to collect national standards, local standards and industry standard files of dangerous chemicals;

[0013] 2.3 Screen and evaluate files: screen and evaluate the collected files, select the files related to the research target; evaluate the credibility and authority of the files;

[0014] 2.4 Summarize and classify files: summarize and classify the screened files, organize the national standards, local standards and industry standard files, and associate the emergency plan and accident disposal specification files with them;

[0015] 2.5 Analyze the content of the files: carefully read and analyze the collected files to understand the specification requirements, safety measures and accident emergency handling process content of dangerous chemicals;

[0016] 2.6 File analysis: Based on the analysis and understanding of the file, write a summary and report of the file, summarize the key points and key information of the domestic standardization files related to dangerous chemicals, including national standards, local standards, industry standards, emergency plans and accident disposal standardization files.

[0017] Further, the collected files are converted into editable text format and formatted, and redundant information is removed to ensure data consistency and availability; the following are the specific tasks and steps:

[0018] 3.1 File format conversion: For the electronic files that have been obtained, check their format and convert them into editable text format;

[0019] 3.2 PDF extraction using OCR technology: Most of the collected electronic files are in PDF format, and the OCR technology is used to extract the text therein to obtain editable text content;

[0020] 3.3 Remove redundant information: After file conversion and OCR extraction, some redundant information exists, and the text is cleaned and formatted using python language to remove redundant information;

[0021] 3.4 Data consistency and proofreading: For the converted text content, proofread and check to ensure data accuracy and consistency; check if the text is correctly extracted;

[0022] 3.5 File naming and organization: name and organize the text files; name the files according to the content, standard number, date factor, and organize and classify them according to certain directory structure.

[0023] The beneficial effects of the present application are that the knowledge graph technology is applied to the field of dangerous chemicals in the petrochemical industry, important information scattered in a large amount of text data is presented and stored in a structured form through standard digitization, which has important significance for the high-quality development of the dangerous chemicals industry. And the present application stores the knowledge of the dangerous chemicals storage area in the petrochemical industry in a more scientific way, improves the knowledge sharing and utilization efficiency in this field. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 is the basic process of constructing the knowledge graph of the dangerous chemicals storage area in the petrochemical industry.

[0025] Figure 2 is a schematic diagram of the semantic relationship between the ontologies of the present application.

[0026] Figure 3 is a schematic diagram of the mode layer of the knowledge graph of the dangerous chemicals storage area in the petrochemical industry in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The application will be further described below with reference to the drawings.

[0028] Please refer to Figure 1 The method constructs a mode layer and a data layer of a knowledge graph based on dangerous chemicals and emergency response files for dangerous chemicals accidents;

[0029] The mode layer, starting from the digitization of dangerous chemical-related standards and emergency response business logic, first finds comprehensive abstract concepts in the field of dangerous chemicals, and then gradually refines upper-level concepts into more specific lower-level concepts from top to bottom. The concept entities, attributes, and hierarchical semantic relationships are defined from top to bottom to construct an accurate and well-structured concept architecture. Through continuous induction, clustering, generalization, and other processing of lower-level concepts and terms in the field knowledge, the abstract concept model of the upper level is synthesized.

[0030] The data layer, from bottom to top, aligns and fuses different data sources such as dangerous chemical-related specification files, academic literature, and accident cases, extracts entity information and semantic associations based on deep learning methods from the characteristics of dangerous chemicals and accident response measures, establishes the mapping between specific elements and concept nodes, forms the mapping from the mode layer to the data layer, and uses a twin neural network-based entity alignment model to fuse multiple knowledge graphs to form a comprehensive dangerous chemical storage area knowledge graph.

[0031] Among them, literature analysis and research are carried out to find the normative files of dangerous chemicals. The following steps are included:

[0032] 2.1 Determine the scope and objectives of the study: read relevant literature to clarify the content and scope of the normative files of dangerous chemicals,

[0033] 2.2 Collect normative files: use academic search engines, library resources, government agency websites, etc. to collect files related to dangerous chemicals, such as national standards, local standards, and industry standards. These documents may include laws and regulations, technical specifications, safety manuals, etc.

[0034] 2.3 Screen and evaluate files: screen and evaluate the collected files, and select the most relevant files to the research objectives. Evaluate the credibility and authority of the files, and preferentially select files published by relevant government agencies, standardization organizations, or authoritative research institutions.

[0035] 2.4 Summarize and classify files: summarize and classify the screened files, organize the relevant files of national standards, local standards, and industry standards, and associate the normative files of emergency plans and accident disposal with them.

[0036] 2.5 Analyze the content of the files: Carefully read and analyze the collected files to understand the regulatory requirements, safety measures, and emergency response procedures for hazardous chemicals. Pay attention to capturing key information such as standard numbers, applicable scope, technical requirements, etc.

[0037] 2.6 File research and analysis: Based on the analysis and understanding of the files, write a summary and report of the domestic regulatory files related to hazardous chemicals, including the key points and information of national standards, local standards, industry standards, emergency plans, and accident disposal. During the entire process, ensure the use of accurate and reliable file sources to ensure the authority and credibility of the information obtained. This will provide important reference for subsequent technical route formulation and hazardous chemical management.

[0038] Convert the collected files into editable text format and perform necessary format conversion, remove redundant information, etc. to ensure data consistency and usability. Here are some specific tasks and steps:

[0039] 3.1 File format conversion: For the electronic files already obtained, check their format and convert them into editable text format. Common file formats include PDF, Word documents, etc.

[0040] 3.2 PDF extraction using OCR technology: Most of the collected electronic files are in PDF format, and the content is scanned images rather than editable text. Use OCR technology to extract the text and obtain editable text content.

[0041] 3.3 Remove redundant information: After file conversion and OCR extraction, there may be some redundant information such as headers, footers, annotations, reference files, etc. Use python language to clean and format the text to remove these redundant information for subsequent data processing and analysis.

[0042] 3.4 Data consistency and proofreading: For the converted text content, conduct a careful proofreading and checking to ensure data accuracy and consistency. Check if the text is correctly extracted, especially for complex content such as numbers, formulas, tables, etc. If necessary, manually correct and adjust the text.

[0043] 3.5 File naming and organization: To facilitate subsequent use and retrieval, it is recommended to name and organize the text files reasonably. Name the files according to their content, standard number, date, etc. and organize and classify them according to certain directory structure.

[0044] The concept hierarchy system of the knowledge graph is designed according to the idea of ontology, arranges the hierarchical relationship, generic relationship and correlation relationship of the concept, and forms a concept framework of the knowledge graph mode layer with clear structure and level. The mode layer ontology library of the knowledge graph contains four core elements of basic attributes, accident events, emergency actions and emergency resources, and defines the semantic relationship between the core elements. The comprehensive ontology library of the knowledge graph mode layer of the dangerous chemical accident emergency can be expressed as:

[0045] O hazardous chemicals ={O BA ,O AE ,O EA ,O ER}

[0046] O BA represents the basic attribute ontology, O AE represents the accident event ontology, O EA represents the emergency action ontology, and O ER represents the emergency resource ontology. Please refer to the description shown in Figure 2 .

[0047] The ontology definitions of the four core elements of basic attributes, accident events, emergency actions and emergency resources in the comprehensive ontology library of the dangerous chemical emergency field are described.

[0048] 4.1 Construction of dangerous chemical basic attribute ontology:

[0049] (1) The basic attribute ontology is a unified description of the concept hierarchical relationship, attribute relationship and correlation relationship of dangerous chemicals. A basic attribute ontology is expressed as:

[0050] O BA ={N C ,N P ,N S ,N D}

[0051] O BA represents the basic attributes of dangerous chemicals, including N C for chemical properties, N P for physical properties, N S for safety and N D for the characteristics of danger;

[0052] Refer to the "Commonly Used Hazardous Chemicals Emergency Handbook", "Hazardous Chemicals Catalogue", "The First Batch of Key Supervision of Hazardous Chemicals Safety Measures and Accident Emergency Disposal Principles" and other national standards. The OBA concept represents the basic properties of hazardous chemicals, including chemical properties (hazardous product name OnA, molecular formula, relative molecular mass, CAS number, harmful substance composition); physical properties (appearance and properties, PH value, melting point, boiling point, relative density, relative vapor density, saturated vapor pressure, critical temperature, critical pressure, LogP); safety and danger (flash point, ignition temperature, combustion heat, lower explosive limit, upper explosive limit, danger class); other characteristics (solubility, main use, stability) and so on.

[0053] 4.2 Construction of hazardous chemical accident case ontology:

[0054] The accident event ontology is a unified description of the hierarchical relationship, attribute relationship and association relationship of the accident event concept. An accident case ontology is represented as:

[0055] O AE ={A FR ,A PS ,A L ,A AL}

[0056] The accident event is divided into two modules of accident type and accident level, in which A FR represents the accident type of fire and explosion accident, A PS represents the accident type of poisoning and suffocation accident, A L represents the accident type of leakage accident, and A AL represents the accident level.

[0057] Construction of emergency action ontology:

[0058] Referring to "AQT 3052-2015 Hazardous Chemicals Emergency Rescue Command Guide", we divide the accident event into two modules of accident type and accident level, in which AFR represents the accident type of fire and explosion accident, APS represents the accident type of poisoning and suffocation accident, AL represents the accident type of leakage accident, and AAL represents the accident level. According to the degree of accident harm and the design range, the accident is divided into four levels: I (particularly serious accident), II (major accident), III (relatively large accident) and IV (general accident).

[0059] Construction of emergency action ontology:

[0060] The emergency action ontology is a unified description of the hierarchical relationship, attribute relationship and association relationship of the emergency action concept. An emergency action ontology is represented as:

[0061] O EA ={M i, M v , M p , M m , M d , M r}

[0062] The emergency action is divided into the following behaviors: M i representing detection, M v representing alert, M p representing protection, M m representing disposal, M d representing decontamination, M r representing recovery.

[0063] Refer to the national standards such as “GAT 970-2011 Hazardous Chemicals Leakage Accident Disposal Action Guidelines”, “XF / T 1275-2015 Oil Storage Tank Fire Fighting Action Guidelines”, “SYT 6306-2008 Normal Pressure Storage Tank Fire Extinguishing Treatment”. The emergency action is divided into the following steps: detection, alert, protection, disposal, decontamination, and recovery.

[0064] Emergency resource ontology construction:

[0065] The emergency resource ontology is a unified description of the concept hierarchical relationship, attribute relationship and association relationship of emergency resources. An emergency resource ontology is represented as:

[0066] O ER ={H pe , H w , H p , H r , H d , H re , H de , H fe , H pd , H ac , H se}

[0067] The emergency resources are divided into: H pe representing protective equipment, H w representing vehicles, H p representing detection equipment, H r representing alert equipment, H d representing communication equipment, H re representing lifesaving equipment, H de representing forcible entry equipment, H fe representing fire extinguishing equipment, H pd representing plugging equipment, H ac representing anti-pollution transportation equipment, and H se representing smoke exhaust and lighting equipment.

[0068] According to the "DB43 / T 1778-2020 Chemical Industry Park Emergency Management and Rescue Specification", emergency resources are divided into: protective equipment, vehicles, detection equipment, warning equipment, communication equipment, lifesaving equipment, forcible entry equipment, fire extinguishing equipment, leak stopping equipment, anti-pollution transportation equipment, smoke exhaust and lighting equipment.

[0069] The processed text needs to be screened, and representative texts are selected for entity, attribute, and relationship labeling. These labels will be used for subsequent BILSTM-CRF model training. At the same time, a supervision and quality control mechanism needs to be established to ensure the accuracy and consistency of the labeling. The following is a detailed description of this step:

[0070] 5.1 Text screening: According to the requirements and goals, determine the screening criteria to select representative samples from the processed text. This can be based on the topic, field, length, etc. of the text.

[0071] 5.2 Labeling entities, attributes, and relationships: Use labeling tools or platforms to manually label the selected text. Determine the entity categories (such as names, places, organizations, etc.), attributes (such as dates, prices, quantities, etc.), and relationships between entities (such as superior-inferior relationships, synonymous relationships, etc.).

[0072] 5.3 Establish a supervision mechanism: It is essential to ensure the accuracy and consistency of the labeling. This can include the following measures:

[0073] 1) Provide labeling guidelines and guidelines: Prepare detailed labeling guidelines and guidelines to ensure that labeling personnel understand the requirements and standards of the labeling task.

[0074] 2) Conduct labeling training: Provide training for labeling personnel to familiarize them with labeling tools and tasks, and to understand labeling guidelines.

[0075] 3) Regularly discuss and solve problems: Hold regular discussions with labeling personnel to answer questions, clarify labeling standards, and promptly address problems encountered during the labeling process.

[0076] 5.4 Establish a quality control mechanism: It is necessary to ensure the consistency and quality of the labeling. The following are some common quality control measures:

[0077] 1) Double labeling: Label a portion of the text by multiple people, then compare the labeling results, evaluate the consistency between labelers, and make necessary corrections and adjustments.

[0078] 2) Regular evaluation and feedback: Regularly evaluate the labeling results, provide feedback to the labeling personnel, and provide guidance to ensure the accuracy and consistency of the labeling.

[0079] 3) Sampling Inspection: Randomly select annotated texts for inspection to verify the accuracy of the annotations and correct any potential errors.

[0080] 5.5 Annotation Tool Selection: Choose an appropriate tool or platform for the annotation task. There are many open-source and commercial annotation tools available, such as Labelbox, Doccano, Brat, etc. Consider factors such as requirements, annotation complexity, team collaboration, and choose the most suitable tool.

[0081] 5.6 Annotation Data Management: Establish an effective annotation data management system to ensure the security and traceability of the annotated data. This can include measures such as data backup, version control, permission management, etc., to prevent data loss or tampering.

[0082] 5.7 Annotation Speed and Quality Balance: Balance the speed and quality of annotation during the annotation process. Try to improve annotation efficiency while ensuring accurate and error-free annotations. This can be achieved by regular checks and feedback, communication with annotators, and providing appropriate time estimates for annotation tasks.

[0083] 5.8 Annotation Result Evaluation: Evaluate and verify the annotation results to ensure accuracy and consistency. This can be done through manual sampling inspection, automatic evaluation indicators, etc. Based on the evaluation results, provide feedback and guidance to annotators and make necessary corrections and adjustments.

[0084] 5.9 Annotation Data Expansion: During the annotation process, consider using annotated samples for model training and expanding annotation data through semi-supervised learning or active learning methods. This helps improve the performance and generalization ability of the model.

[0085] 5.10 Data Privacy and Security: When conducting manual annotation, ensure proper handling of sensitive information and take appropriate data security measures. This includes limiting access permissions for annotators, anonymizing data, etc., to protect data privacy and security.

[0086] The BILSTM-CRF (Bidirectional Long Short-Term Memory Network-Conditional Random Field) model will be used for named entity recognition. This model combines BILSTM and CRF two parts, which can effectively capture the context information in the text and the dependency relationship between labels, so as to achieve accurate entity recognition. The following is a detailed description of this step:

[0087] 6.1 Data Preparation: Prepare the annotated data set, which contains texts that have been manually annotated and corresponding entity labels. Ensure the quality of the data set and the accuracy of the annotations.

[0088] 6.2 Data Preprocessing: Preprocess the data, including tokenization, constructing word vectors, converting text to a format acceptable by the model, etc.

[0089] 6.3 Feature Extraction: Extract features from the preprocessed data for use in the BILSTM-CRF model. Commonly used features include word vectors, character-level features, part-of-speech tags, etc. Choose appropriate features based on actual needs.

[0090] 6.4 Model Construction: Construct the BILSTM-CRF model. This model usually consists of two parts: a BILSTM layer to capture contextual information and a CRF layer to model the transition probabilities between labels. Use deep learning frameworks such as TensorFlow, PyTorch, or Keras to implement the model.

[0091] 6.5 Model Training: Train the BILSTM-CRF model using the annotated dataset. Divide the dataset into training, validation, and test sets, and iteratively optimize model parameters to improve model performance. Choose appropriate loss functions (such as cross-entropy loss function) and optimization algorithms (such as Adam or SGD) for model training.

[0092] 6.6 Hyperparameter Tuning: Adjust the hyperparameters of the model to achieve better performance. Hyperparameters include learning rate, hidden layer size, dropout rate, etc. Use techniques such as cross-validation to find the best combination of hyperparameters.

[0093] 6.7 Model Evaluation: Evaluate the trained BILSTM-CRF model using the test set. Common evaluation metrics include accuracy, recall, F1 value, etc. Measure the performance of the model through evaluation results and make necessary adjustments and improvements.

[0094] 6.8 Model Application: Use the trained BILSTM-CRF model for named entity recognition in practical applications. Input the text to be recognized into the model and obtain the predicted entity labels

[0095] 6.9 Error Analysis and Tuning: Perform error analysis on the model to identify common error types and patterns. For example, the model may not accurately identify certain types of entities or specific contextual situations. Based on the results of error analysis, adjust the model's architecture, feature selection, or data preprocessing methods to improve the performance of the model.

[0096] 6.10 Model Optimization and Iteration: Based on the results of error analysis and tuning, optimize the model and perform iterative training. May need to try different model architectures, feature engineering methods, or hyperparameter settings to further improve the accuracy and robustness of named entity recognition.

[0097] 6.11 Model Deployment: Deploy the trained BILSTM-CRF model into a real production environment. Integrate the model into an application or system, and ensure it can efficiently process input text and output accurate named entity recognition results.

[0098] 6.12 Continuous Improvement and Updating: Continuously monitor the performance of the model and make improvements and updates based on feedback and needs in actual applications. Consider retraining the model periodically, using larger-scale datasets, or introducing other techniques and models to further improve the effectiveness of named entity recognition.

[0099] The method further comprises constructing a petrochemical industry dangerous chemical warehouse area database based on a Neo4j graph database foundation, to realize dangerous chemical warehouse area knowledge graph storage and dangerous chemical warehouse area visualization display.

[0100] Neo4j is a widely used graph database suitable for building knowledge graphs. It provides powerful graph database functions and Cypher query language, which can effectively store and query graph data. The following are the detailed steps of building a knowledge graph using Neo4j:

[0101] 7.1 Define entity and relationship types: According to the requirements, determine the entity types and relationship types in the knowledge graph. Create corresponding node labels and relationship types to store data into Neo4j.

[0102] 7.2 Import data: Import entity and relationship data into Neo4j. You can use Cypher query language or import tools provided by Neo4j (such as LOAD CSV) to import data from external sources into the database. Ensure that the data format meets the requirements of Neo4j, and perform necessary data cleaning and conversion.

[0103] 7.3 Create nodes and relationships: Use Cypher language to create nodes and relationships. By using the CREATE statement, create a node for each entity, and use the relationship statement to create relationships between entities. Set the properties of nodes and relationships to add relevant information.

[0104] 7.4 Query and retrieval: Use Cypher query language to query and retrieve graph data. According to the requirements, write corresponding query statements to obtain specific entity, relationship or attribute information. You can use MATCH, CREATE, DELETE and other statements to manipulate data.

[0105] 7.5 Indexing and Optimization: Create indexes for graph data to improve query performance. Neo4j supports creating indexes on nodes and properties to speed up queries. Based on the frequency and performance requirements of queries, choose appropriate properties to index and use the PROFILE command to analyze query performance.

[0106] 7.6 Updating and Maintenance: Update and maintain the knowledge graph as needed. Add new entities and relationships, update existing properties or relationships, and ensure the knowledge graph stays synchronized with actual data. Perform necessary data cleaning and data quality control to ensure the accuracy and consistency of the graph.

[0107] 7.7 Visualization and Exploration: Visualize and explore the knowledge graph using visualization tools provided by Neo4j, such as Neo4j Browser, Neo4j Bloom, or other third-party tools. Through graphical interfaces, visually browse and understand entities, relationships, and properties in the graph.

[0108] By following the above steps, you can build a functional knowledge graph using Neo4j. Utilize the graph database features and powerful query language of Neo4j to perform complex graph data operations and efficient knowledge discovery. Remember to index and optimize performance as needed during the construction process to achieve better query efficiency and user experience.

[0109] The ontology of the dangerous chemical warehouse area knowledge graph takes dangerous chemicals in the warehouse area and emergency response to dangerous chemical accidents as the core (see Figure 3 The dangerous chemical-related knowledge mainly revolves around its own physicochemical properties and storage conditions in the warehouse area. The important node of the dangerous chemical accident in the warehouse area is emergency action, which connects emergency rescue actions and emergency resources through emergency action.

[0110] The dangerous chemical warehouse area knowledge graph is based on the type and physicochemical properties of dangerous chemicals in the petrochemical industry. Different accident types can lead to different accident consequences, and targeted emergency measures need to be taken at each link of emergency rescue, and appropriate emergency resources need to be called.

[0111] The patent applies knowledge graph technology to the field of dangerous chemicals in the petrochemical industry, and through standard digital means, important information scattered in a large amount of text data is presented and stored in a structured form, which is of great significance to the high-quality development of the digitalization of the dangerous chemical industry. The patent stores the knowledge of the dangerous chemical storage area in the petrochemical industry in a more scientific way, and improves the knowledge sharing and utilization efficiency in this field. The beneficial effect of the patent is that in a specific dangerous chemical accident situation, real-time emergency resource demand will be generated, the knowledge graph technology is used, the properties and storage conditions of dangerous chemicals in the storage area are stored, and the emergency resources required under different emergency rescue actions after different types of dangerous chemical accidents caused by these dangerous chemicals. Provide a strong reference for the decision-making of dangerous chemical accident emergencies to support dangerous chemical accident emergency resource information visualization and emergency response scheme reasoning, improve emergency response speed and decision-making efficiency, and promote the intelligentization and precision of emergency management work.

[0112] The above only describes the preferred embodiments of the present application, and any equivalent changes and modifications made within the scope of the patent application of the present application shall be within the scope of the present application.

Claims

1. A method for constructing a knowledge graph of a petrochemical industry dangerous chemical warehouse area based on a document, characterized in that: The method constructs a mode layer and a data layer of a knowledge graph of dangerous chemicals and emergency disposal files of dangerous chemical accidents; The mode layer defines concept entities, attributes and hierarchical semantic relationships from top to bottom, and constructs an accurate and structured hierarchical concept architecture; The data layer aligns and fuses different source knowledge from bottom to top according to the characteristics of dangerous chemicals and the measures taken in the accident response based on deep learning to extract entity information and semantic association, establish the mapping between specific elements and concept nodes, form the mapping from the mode layer to the data layer, and use the entity alignment model based on the twin neural network to fuse multiple knowledge graphs to form a comprehensive dangerous chemical storage knowledge graph; The mode layer ontology library of the knowledge graph includes four core elements of basic attributes, accident events, emergency actions and emergency resources, and the comprehensive ontology library of the mode layer of the dangerous chemical accident emergency knowledge graph is represented as: O hazardouschemicals = {O BA , O AE , O EA , O ER} O BA represents a basic property ontology, O AE represents an incident event ontology, O EA represents an emergency action ontology, O ER represents an emergency resource ontology; A basic attribute ontology is a unified description of the hierarchical relationship, attribute relationship and association relationship of dangerous chemical concepts, and is represented as: O BA = {N C , N P , N S , N D} O BA The meaning represents the basic attribute of the dangerous chemical, including N C Chemical properties, N P Physical properties, N S Safety and N D The characteristics of the danger; An accident event ontology is a unified description of the hierarchical relationship, attribute relationship and association relationship of accident event concepts, and is represented as: O AE = {A FR , A PS , A L , A AL} The accident event is divided into two modules of accident type and accident level, wherein A FR represents the accident type of fire and explosion accident, A PS represents the accident type of poisoning and suffocation accident, A L represents the accident type of leakage accident, A AL represents the accident level; the emergency action ontology is a unified description about the concept hierarchical relationship, attribute relationship and association relationship of the emergency action, and an emergency action ontology is represented as: O EA = {M i , M v , M p , M m , M d , M r} M i representing detection, M v representing alert, M p representing protection, M m representing disposal, M d representing decontamination, M r representing recovery; the emergency resource ontology is a unified description of the hierarchical relationship, attribute relationship and association relationship of the emergency resource concept, and an emergency resource ontology is represented as: O ER = {H pe , H w , H p , H r , H d , H re , H de , H fe , H pd , H ac , H se} H pe representing protective equipment, H w representing vehicles, H p representing detection equipment, H r representing alerting equipment, H d representing communication equipment, H re representing lifesaving equipment, H de representing forcible entry equipment, H fe representing fire extinguishing equipment, H pd representing plugging equipment, H ac representing anti-pollution transport equipment, H se representing smoke evacuation lighting equipment.

2. The method according to claim 1, wherein the method is characterized by: Academic literature analysis and research are conducted to find the specification files of dangerous chemicals, including: Determine the scope and goal of the study: read academic literature to clearly understand the content and scope of the specification files of dangerous chemicals; Collect specification files: use academic search engines, library resources, and government agency website channels to collect national, local and industry standard files related to dangerous chemicals; Screen and evaluate files: screen and evaluate the collected files, select files related to the research goal, and evaluate the credibility and authority of the files; Summarize and classify files: summarize and classify the screened files, organize national, local and industry standard files, and associate emergency plans and accident disposal specification files; Analyze file content: carefully read and analyze the collected files to understand the specification requirements, safety measures and accident emergency handling process content of dangerous chemicals; File research: based on the analysis and understanding of the files, write a file review and report to summarize the key points and critical information of the specification files of dangerous chemicals in China, including national, local, industry standards, emergency plans and accident disposal specification files.

3. The method according to claim 1, wherein the method is characterized by: Convert the collected files into editable text format and perform format conversion and redundant information removal operations to ensure data consistency and usability; including: File format conversion: check the format of the obtained electronic files and convert them into editable text format; PDF uses OCR technology to extract: most of the collected electronic files are in PDF format, and the OCR technology is used to extract the text to obtain editable text content; Remove redundant information: after file conversion and OCR extraction, some redundant information exists, and python language is used to clean and format the text to remove redundant information; Data consistency and proofreading: proofread and check the converted text content to ensure data accuracy and consistency; check if the text is correctly extracted; File naming and organization: Name and organize text files according to their content, standard number, date factors, and organize them in a certain directory structure.

4. The method according to claim 3, wherein the method is characterized by: Filter the processed text and select representative text for entity, attribute, and relationship annotation; These annotations will be used for subsequent BILSTM-CRF model training, while a supervision and quality control mechanism will be established to ensure the accuracy and consistency of the annotations; including: Text filtering: Determine the filtering criteria according to the needs and goals to select representative samples from the processed text; Annotate entities, attributes, and relationships: Use annotation tools or platforms to manually annotate selected text; determine the entity categories, attributes, and relationships between entities; Establish a supervision mechanism: Ensure the accuracy and consistency of the annotations, establish a supervision mechanism, including: Provide annotation guidelines and guidelines: Prepare annotation guidelines and guidelines to ensure that annotators understand the requirements and standards of the annotation task; Provide training for annotators: Familiarize annotators with annotation tools and tasks, and understand annotation guidelines; Regularly discuss and solve problems: Regularly discuss with annotators, answer questions, clarify annotation standards, and timely solve problems encountered in the annotation process; Establish a quality control mechanism: Ensure the consistency and quality of the annotations, establish a quality control mechanism, including: Double annotation: A portion of the text is annotated by multiple people, then the annotation results are compared, the consistency between annotators is evaluated, and corrections and adjustments are made; Regular evaluation and feedback: Regularly evaluate the annotation results, provide feedback to the annotators, and provide guidance to ensure the accuracy and consistency of the annotations; Sample inspection: Randomly select annotated text for inspection to verify the accuracy of the annotations and correct any errors; Annotation tool selection: Select the tool or platform for the annotation task; Tools include: Labelbox tool, Doccano tool, Brat tool; Annotation data management: Establish an effective annotation data management system to ensure the security and traceability of the annotation data; Implement data backup, version control, and permission management measures to prevent data loss or tampering; Balance annotation speed and quality: Balance the annotation speed and quality during the annotation process; Achieve this by regularly checking and providing feedback to annotators, as well as providing time estimates for the annotation task; Annotation result evaluation: Evaluate and verify the annotation results to ensure the accuracy and consistency of the annotations; Use manual sampling inspection and automatic evaluation indicators; Based on the evaluation results, provide feedback and guidance to the annotators, and make corrections and adjustments; Annotation data expansion: During the annotation process, consider using the annotated samples for model training, and expand the annotation data through semi-supervised learning or active learning; Data privacy and security: When performing manual annotation, ensure that sensitive information is properly handled and take data security protection measures; This includes limiting the access rights of annotators, anonymizing data to protect the privacy and security of data.

5. The method according to claim 4, wherein the method is characterized by: Entity recognition will be performed using the BILSTM-CRF model; this model combines a bidirectional long short-term memory network (BILSTM) and a conditional random field (CRF) to effectively capture contextual information and dependencies between labels, enabling accurate entity recognition; including: Data preparation: Prepare a labeled dataset containing manually annotated text and corresponding entity labels; ensure the quality and accuracy of the dataset; Data preprocessing: Preprocess the data, including tokenization, word vector construction, and converting text into a format that the model can accept; Feature extraction: Extract features from preprocessed data for use by the BILSTM-CRF model; features include word vectors, character-level features, and part-of-speech tags; Model construction: Build the BILSTM-CRF model; this model consists of two parts: a BILSTM layer to capture contextual information and a CRF layer to model the transition probabilities between labels; use a deep learning framework to implement the model; the deep learning framework includes TensorFlow, PyTorch, or Keras; Model training: Train the BILSTM-CRF model using the labeled dataset; divide the dataset into training, validation, and test sets, and iteratively optimize model parameters to improve model performance; select a loss function and optimization algorithm for model training; Hyperparameter tuning: Adjust the model's hyperparameters to achieve better performance; hyperparameters include learning rate, hidden layer size, and dropout rate; use cross-validation techniques to find the optimal combination of hyperparameters; Model evaluation: Evaluate the trained BILSTM-CRF model using the test set; evaluation metrics include accuracy, recall, and F1 score; use the evaluation results to measure the model's performance and make adjustments and improvements; Model application: Use the trained BILSTM-CRF model for entity recognition in practical applications; input the text to be recognized into the model and obtain the predicted entity labels; Error analysis and tuning: Perform error analysis on the model to identify error types and patterns; i.e., the model does not accurately identify certain types of entities or contextual situations; based on the error analysis results, adjust the model's architecture, feature selection, or data preprocessing methods to improve the model's performance; Model optimization and iteration: Based on the results of error analysis and tuning, optimize the model and perform iterative training; try different model architectures, feature engineering methods, or hyperparameter settings to improve the accuracy and robustness of entity recognition; Model deployment: Deploy the trained BILSTM-CRF model to a production environment; integrate the model into an application or system and ensure it can efficiently process input text and output accurate entity recognition results; continuously improve and update: continuously monitor the model's performance and make improvements and updates based on feedback and requirements in practical applications.

6. The method according to claim 1, wherein the method is characterized by: The method also includes constructing a petrochemical industry dangerous chemical warehouse area database based on a Neo4j graph database foundation based on the dangerous chemical warehouse area knowledge graph, to realize dangerous chemical warehouse area knowledge graph storage and dangerous chemical warehouse area visualization display; including: Defining entity and relationship types: determine entity types and relationship types in the knowledge graph according to requirements; create corresponding node labels and relationship types to store data into Neo4j; Import data: import entity and relationship data into Neo4j, use Cypher query language or import tools provided by Neo4j to import data from external sources into the database; Create nodes and relationships: use Cypher language to create nodes and relationships, create a node for each entity by using the CREATE statement, and use the relationship statement to create relationships between entities; set node and relationship properties to add corresponding information; Query and retrieval: use Cypher query language to query and retrieve graph data; write corresponding query statements according to requirements to obtain entity, relationship or attribute information; Index and optimization: create indexes for graph data to improve query performance; Neo4j supports creating indexes for nodes and attributes to speed up queries; according to the frequency and performance requirements of queries, select the corresponding attributes for indexing, and use the PROFILE command to analyze query performance; Update and maintenance: update and maintain the knowledge graph as needed; add new entities and relationships, update existing attributes or relationships, and ensure that the knowledge graph is synchronized with actual data; Visualization and exploration: use the visualization tools provided by Neo4j to visualize and explore the knowledge graph; through the graphical interface, intuitively browse and understand the entities, relationships and attributes in the graph.

Citation Information

Patent Citations

  • Industry process field knowledge graph construction method and device

    CN111444351A

  • Dangerous chemical accident knowledge base construction method based on knowledge graph

    CN115953117A