A standard metadata storage method and machine-readable storage medium

By constructing a knowledge graph of coal chemical industry topology and dynamic coding, the problem of storing and managing multimodal data in the coal chemical industry has been solved. This has enabled cross-process unit data analysis and fault tracing, optimized resource allocation and data quality, and improved storage and query efficiency.

CN121436137BActive Publication Date: 2026-04-14CHINA NAT INST OF STANDARDIZATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In the coal chemical industry, the storage and management of multimodal data suffers from data silos. Traditional storage solutions are difficult to adapt to dynamic process changes, resulting in low data query efficiency, resource waste, and inconsistent data quality, which cannot support data analysis and fault tracing across process units.

Method used

By collecting and preprocessing coal chemical operation data and related data, multimodal semantic information is extracted and structured, a coal chemical topology knowledge graph is constructed and dynamically encoded, a standard metadata storage model is established, and a three-level verification method is used to resolve data contradictions, thereby achieving cross-modal data fusion and full lifecycle management.

Benefits of technology

It enables cross-process unit data analysis and fault tracing, dynamically adapts to process changes, optimizes resource allocation and data quality, improves storage and query efficiency, and supports process optimization and fault tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436137B_ABST
    Figure CN121436137B_ABST
Patent Text Reader

Abstract

The application discloses a standard metadata storage method and a machine readable storage medium, comprising collecting operation data and associated data of a preset coal chemical industry, and preprocessing the equipment data and the associated data; performing multi-modal semantic information extraction on the associated data to obtain information data, performing structured processing on the information data and the operation data to obtain structure data, and constructing a process mapping table according to business terms and the process parameters; the information data comprises incremental information and general information; constructing a coal chemical industry topology knowledge graph according to the structure data and the process mapping table based on a chemical industry association type, performing time sequence coding on the coal chemical industry topology knowledge graph according to a time window and a process unit to obtain dynamic coding; and constructing a coal chemical industry standard metadata storage model according to the dynamic coding, inputting to-be-stored data into the coal chemical industry standard metadata storage model, and outputting a storage result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a standard metadata storage method and a machine-readable storage medium. Background Technology

[0002] In the process of digital transformation in the coal chemical industry, the entire production process involves operational data such as process parameters, equipment static attributes, and material and energy data, as well as various types of related data such as design documents, equipment images, and monitoring videos. The data sources are scattered and the formats are heterogeneous. Traditional metadata storage methods lack a unified structured processing mechanism, making it difficult to achieve effective association of multimodal data, resulting in prominent data silo problems and an inability to support data analysis needs across process units.

[0003] Existing storage solutions are not sufficiently adaptable to dynamically changing process parameters, and the process mapping table is updated lagging behind, making it difficult to meet the dynamic needs of coal chemical process optimization and standard specification iteration. Furthermore, data storage does not employ differentiated management based on process logic and access frequency, resulting in low query efficiency for high-frequency data and excessive storage resource consumption for low-frequency data, thus impacting data utilization efficiency.

[0004] Furthermore, information inconsistencies easily arise during multi-source data fusion, and the lack of a scientific verification mechanism leads to inconsistent metadata quality. Traditional methods have not constructed an effective data lineage tracing and topological association system, failing to achieve visualized management of the entire data lifecycle, thus hindering in-depth application in business scenarios such as process optimization and fault tracing. Therefore, there is an urgent need for a standard metadata storage method that adapts to the characteristics of the coal chemical industry, supports multimodal data fusion, and possesses dynamic adaptability and efficient storage capabilities. Summary of the Invention

[0005] The purpose of this invention is to provide a standard metadata storage method and a machine-readable storage medium.

[0006] To achieve the above objectives, the present invention is implemented according to the following technical solution:

[0007] This invention includes the following steps:

[0008] Collect pre-defined operational data and related data for coal chemical engineering, and preprocess the equipment data and related data; the operational data includes process parameters, process condition data, equipment static attributes, and material and energy data; the related data includes document data and multimedia data; the document data includes design documents, analysis reports, and compliance documents; the multimedia data includes equipment images and video streams;

[0009] Multimodal semantic information extraction is performed on the associated data to obtain information data; the information data and the operation data are subjected to structured processing to obtain structured data; and a process mapping table is constructed based on business terms and process parameters; the information data includes incremental information and general information.

[0010] Based on the chemical industry association type, a coal chemical industry topology knowledge graph is constructed according to the structural data and the process mapping table. The coal chemical industry topology knowledge graph is then time-series encoded according to the time window and process unit to obtain dynamic encoding.

[0011] A coal chemical industry standard metadata storage model is constructed based on the dynamic encoding. The data to be stored is input into the coal chemical industry standard metadata storage model, and the storage result is output.

[0012] Furthermore, the method for extracting information data from the associated data using multimodal semantic information includes:

[0013] Extract general information: Extract process-related information from associated data to obtain equipment performance characteristics, material conversion patterns, and energy consumption coupling relationships; perform semantic annotation on document data to extract explicit and implicit knowledge to obtain key points of process procedures, analysis report conclusions, and design specification requirements; perform visual semantic extraction on multimedia data to transform it into describable information to obtain equipment status indicators, abnormal operating condition indicators, and environmental feature extraction.

[0014] Incremental information extraction: Extract equipment tag number, medium name, and operating pressure triplets from flow charts; extract sample number, test items, values, and compliance standards from analysis reports; extract step numbers, operating actions, and process parameters from operating procedures; directly extract incremental information from document data; for multimedia data: extract location, corrosion area, and severity from reactor lining corrosion photos in equipment status diagrams; perform OCR recognition on instrument reading images to obtain pressure gauge and thermometer readings and associate them with corresponding measurement point codes; extract keyframes from tank area monitoring videos to obtain liquid level, time, and abnormal status; identify maintenance operation videos to obtain tool type, operating steps, and personnel qualifications;

[0015] Using equipment, materials, processes, and environment from general information as core entities, establish multi-dimensional relationships; perform semantic association based on incremental information to obtain image-real-time data association and video-fault tracing association;

[0016] When multi-source data encounters information contradictions, a three-level verification method is used for verification: priority determination, time correlation, and credibility weighting. Priority determination is: online analytical instrument data > laboratory test reports > manual records. Time correlation prioritizes multimodal data within the same time window. Credibility weighting assigns credibility coefficients to data from different sources, and the weighted values ​​are used for fusion when conflicts occur.

[0017] The multi-dimensional correlation, image-real-time data correlation, and video-fault tracing correlation are output as information data.

[0018] Furthermore, a method for obtaining structured data by structuring the information data and the job data includes:

[0019] For information data: Through three steps of extraction, cleaning, and association, key information is extracted, missing values ​​are handled, and outliers are corrected. Based on the information data, a chain traceability is achieved through foreign key associations to realize raw coal batches → quality inspection reports → supplier information. Relationship modeling is performed based on the chain traceability to obtain structured information data.

[0020] A three-level coding rule is used to assign a unique identifier to each monitoring point in the operation data; the five elements of the operation data, namely equipment number, data category, timestamp, value, and quality mark, are recorded according to the monitoring point to obtain structured operation data; the database is partitioned by process unit and day to obtain hot data and cold data; the hot data area stores the current and the most recent day's second / minute level data, the warm data area stores downsampled data from 1 to 30 days, and the cold data area stores archived data from more than 30 days.

[0021] Attach association pointers to unstructured files, obtain the relationship between equipment, materials, and parameters based on structured information data and structured operational data, construct a structured knowledge graph based on the relationship between equipment, materials, and parameters, record the entire data lifecycle, and obtain cross-modal data associations;

[0022] Output the structured knowledge graph as structured data.

[0023] Furthermore, the method for constructing a process mapping table based on business terms and the process parameters includes:

[0024] Based on business terminology, key parameters are extracted from the process parameters. These key parameters include process unit, process parameter, number type, unit, value range, associated equipment ID, and data source system. A process mapping table is then constructed based on the business terminology and key parameters.

[0025] Establish a dynamic update mechanism. After accepting an update request, modify the process mapping table in the data management platform and update the process unit, process parameters, number type, unit, value range, associated equipment ID, and data source system synchronously. After the update, verify the consistency of the updated data. The process mapping table adopts the naming rule of basic version and change number. Each update records the person making the change, the time of the change, and the content of the change.

[0026] When process parameters change, business terms are added, or standards and specifications are updated, a dynamic update mechanism is triggered to update the process mapping table and output the updated process mapping table.

[0027] Furthermore, methods for constructing a knowledge graph of coal chemical industry topology include:

[0028] With process flow, material flow, and energy flow as the core, the process units, equipment, materials, and parameter elements of the entire coal chemical process are abstracted into nodes through entity and relation modeling, and directed edges are used to represent the logical relationships between elements.

[0029] Entity types are categorized based on the characteristics of structured data and process logic. Each entity type contains standardized attributes, including process units, equipment, materials, process parameters, operating conditions, and analysis indicators.

[0030] Based on the definition of chemical process logic, the semantic relationship between source entities and target entities is clarified. Combining the characteristics of coal chemical process, core association types are extracted, including inclusion, feeding, generation, control, consumption, satisfaction, association, and flow direction. The direction of inclusion is from process unit to equipment; the direction of feeding is from material to equipment; the direction of generation is from equipment to material; the direction of control is from process parameters to equipment; the direction of consumption is from equipment to material; the direction of satisfaction is from equipment to operating conditions; the direction of association is from analytical indicators to material; and the direction of flow direction is from material to process unit.

[0031] The knowledge graph core framework is constructed based on entity type and core association type. The topology structure is decomposed according to the five major process units of raw material pretreatment, gasification, purification, synthesis and separation. The connection logic of elements within and between units is clarified to obtain the core process topology structure of coal chemical industry.

[0032] Construct the topology of the raw material pretreatment unit, and obtain the core entities and relationships based on the entity types of the raw material pretreatment unit. The entity types include process unit, equipment, material and process parameters. The core entity of the process unit is the raw material pretreatment unit, the core entities of the equipment are crushing equipment and dryer, the core entities of the material are raw coal and crushed coal, and the core entities of the process parameters are crushing particle size and drying temperature.

[0033] Construct the topology of the gasification unit, and obtain the core entities and relationships through the entity types of the gasification unit. The entity types include process unit, equipment, material and operation unit. The core entity of the process unit is the gasification unit, the core entities of the equipment are the gasifier and burner, the core entity of the material is the crude syngas, and the core entities of the process parameters are the gasifier temperature and oxygen-coal ratio.

[0034] Construct the purification unit topology, and obtain the core entities and relationships through the entity types of the purification units. The entity types include process units, equipment, materials and operation units. The core entity of the process unit is the purification unit, the core entity of the equipment is the shift reactor, the core entity of the material is the purified gas, and the core entity of the process parameters is the catalyst activity conditions.

[0035] Construct the topology of synthesis and separation units, and obtain the core entities and relationships based on the entity types of synthesis and separation units. Entity types include process units, equipment, materials and analytical indicators. The core entities of process units are synthesis units and separation units, the core entities of equipment are synthesis tower and ethylene distillation tower, the core entity of materials is polymer-grade ethylene, and the core entity of analytical indicators is ethylene purity.

[0036] Based on the core framework of the knowledge graph and the core process topology of coal chemical industry, a knowledge graph of coal chemical industry topology structure is output.

[0037] Furthermore, the method for obtaining dynamic coding by performing time-series coding on the coal chemical industry topology knowledge graph includes:

[0038] Historical data is acquired to construct a static coding baseline. A hierarchical dynamic coding system is then employed, fusing process unit attributes, time window features, and data popularity weights. Priority is dynamically adjusted using a popularity evaluation function, the expression of which is:

[0039]

[0040] in For heat evaluation function, This is the process correlation coefficient. This represents the current data query frequency. This represents the system's maximum query frequency. The standard deviation of the data. The maximum allowable standard deviation of process variation. The weighting coefficient for access frequency. This is a weighting coefficient for process correlation. The weighting coefficient for volatility sensitivity;

[0041] Historical data is used to train the process correlation coefficient and optimize parameters. Given a priority dynamic adjustment rule, when the heat evaluation value is greater than or equal to 0.802, the encoding priority is the highest, and an independent memory cache is allocated; when the heat evaluation value is greater than or equal to 0.51 and less than 0.802, the priority is medium, and the data is input into the hot data pool and updated hourly; when the heat evaluation value is less than 0.51, the priority is low, and the data is input into the cold data archive.

[0042] A time window sliding mechanism is set up: a short window is used to store high-frequency data at the second level, and the complete timestamp is retained during encoding; a medium window is used to downsample to the minute-level mean, and similar detection points are merged during encoding; a long window is used to compress hour-level feature values, and only key indicators are retained during encoding.

[0043] Deploy time window sliding triggers based on the time window sliding mechanism, and perform cross-unit data association through process knowledge graph edge attribute encoding. The edge attribute fields are composed of source unit code, target unit code, association type code and time delay in sequence, and the hierarchical dynamic encoding is output as dynamic encoding.

[0044] Furthermore, the method for constructing a coal chemical industry standard metadata storage model based on the dynamic encoding includes:

[0045] The coal chemical industry standard metadata storage model comprises four modules: perception, storage, decision-making, and application. It achieves dynamic metadata growth through an algorithm chain. The perception layer deploys entity recognition and natural language processing algorithms to extract entities, attributes, and relationships from multi-source heterogeneous data sources, forming initial triples. The storage layer uses the Neo4j graph database to store the knowledge graph, a time-series database to store dynamic metadata, and object storage to associate unstructured files, achieving cross-database association through unified encoding. The decision-making layer runs anomaly detection and clustering / classification algorithms to verify metadata quality in real time and dynamically optimize the knowledge graph structure. The application layer, based on similarity search and link prediction algorithms, outputs business services and feeds back feedback data to the perception layer, forming a self-iteratory closed loop.

[0046] Develop dynamic coding rules and attach lifecycle status codes and association strength weights to each metadata element. The dynamic coding rules include four layers: metadata collection and construction, metadata storage and indexing, metadata quality and intelligent management, and metadata application and value mining.

[0047] Metadata collection and construction: A hybrid strategy of rule-based and deep learning is adopted to address the multi-source nature of equipment names. Rule cleaning is performed based on an industry terminology database, string similarity is calculated and embedded into a knowledge graph. When a new entity is added to the database, its contextual relationship is learned through node embedding vectors. After comparison with existing entity vectors, it is automatically classified. Industrial fine-tuning is performed based on the BERT-base model to identify equipment name, parameter, and unit triplet. Few-shot learning is used to process scarce labeled data. Relationships are identified through trigger words and dependency syntax, and an asset hierarchy tree is automatically constructed.

[0048] Metadata Storage and Indexing: Neo4j graph database is used as the core storage engine. Graph algorithms enable efficient organization and dynamic querying of metadata relationships. Key encoding strategies include dynamic graph structure optimization, community detection, and dynamic clustering. Time-series metadata is stored in partitions by device ID and time window. Through APOC plugin linkage, millisecond-level relational queries of real-time data, static attributes, and relationship networks are achieved. Dynamic graph structure optimization involves hierarchical storage of core entities, related entities, and attribute entities. Core entities use memory-first indexing, and weight values ​​are assigned to directed edges to support Dijkstra's algorithm for finding critical dependency paths. Community detection and dynamic clustering identify tight clusters in the metadata graph through community detection algorithms and automatically generate process unit knowledge packages. The weight values ​​assigned to directed edges are calculated based on a comprehensive consideration of process association strength, data flow frequency, and business importance.

[0049] Metadata Quality and Intelligent Management: A three-layer defense is deployed to address common metadata quality issues, identifying isolated entities and automatically triggering review when the anomaly score > 0.85; GNN-based reconstruction of normal connection patterns with an error > 0.3 is used to classify anomalies; historical metadata change records are compared, and if the timestamp deviation from the process change order is > 24 hours, it is automatically marked as pending synchronization; for new metadata, semi-supervised learning is used for automatic classification, dividing new measurement points into multiple initial clusters based on collection frequency, engineering unit, and region, and training a model based on existing labels; when metadata is updated, it is automatically pushed to multiple related systems via graph query statements, with a synchronization delay ≤ 15 minutes;

[0050] The metadata application and value mining process consists of four layers: comparing equipment type, process location, and key parameters for attribute matching; using graph neural networks to aggregate neighbor information and output a comprehensive similarity score; learning the embedded representation of known metadata relationships through graph neural networks; and using node similarity and structural features, where metadata relationships include causal relationships and process optimization.

[0051] Secondly, embodiments of this application also provide an electronic device, including:

[0052] A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the method described in the first aspect.

[0053] Thirdly, embodiments of this application also provide a computer-readable storage medium storing one or more programs that, when executed by an electronic device including multiple applications, cause the electronic device to perform the steps of the method described in the first aspect.

[0054] The beneficial effects of this invention are:

[0055] This invention provides a standard metadata storage method and a machine-readable storage medium. Compared with existing technologies, this invention has the following technical advantages:

[0056] This invention addresses the issue of isolated multi-source data by employing preprocessing, multimodal semantic information extraction, structured processing, construction of a process mapping table, construction of a coal chemical industry topology knowledge graph, dynamic coding, and model building steps. Through multimodal semantic extraction and structured processing, it connects operational data with documents and multimedia data, achieving cross-modal data fusion and supporting cross-process unit analysis. It dynamically adapts to process changes, with a dynamic update mechanism for the process mapping table responding to parameter changes and terminology additions while ensuring data consistency. It improves storage and query efficiency by using layered encoding and storage based on popularity and time windows, prioritizing the caching of high-frequency data and archiving low-frequency data, thus optimizing resource allocation. It ensures data quality and traceability by employing a three-level verification method to resolve data inconsistencies, and a topology knowledge graph to achieve full lifecycle data management, facilitating process optimization and fault tracing. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating the steps of a standard metadata storage method according to the present invention;

[0058] Figure 2 This is a schematic diagram of the structure of an electronic device in an embodiment of this specification. Detailed Implementation

[0059] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.

[0060] The present invention discloses a standard metadata storage method and a machine-readable storage medium method, comprising the following steps:

[0061] like Figure 1 As shown, this embodiment includes the following steps:

[0062] Collect pre-defined operational data and related data for coal chemical engineering, and preprocess the equipment data and related data; the operational data includes process parameters, process condition data, equipment static attributes, and material and energy data; the related data includes document data and multimedia data; the document data includes design documents, analysis reports, and compliance documents; the multimedia data includes equipment images and video streams;

[0063] In actual assessments, operational data includes process parameters such as gasifier temperature (800-1400℃) and oxygen-to-coal ratio (0.8-1.2), static properties such as equipment material and rated pressure, and material and energy data such as raw coal consumption and steam production. Related data includes document data such as process flow diagrams and catalyst analysis reports, and multimedia data such as reactor lining photos and tank area monitoring videos.

[0064] Standardize the format of equipment data and unify the units of parameters exported from different systems; deduplicate document data and remove duplicate design materials; convert the format of multimedia data, unifying video streams to MP4 format and images to JPG format.

[0065] Multimodal semantic information extraction is performed on the associated data to obtain information data; the information data and the operation data are subjected to structured processing to obtain structured data; and a process mapping table is constructed based on business terms and process parameters; the information data includes incremental information and general information.

[0066] Based on the chemical industry association type, a coal chemical industry topology knowledge graph is constructed according to the structural data and the process mapping table. The coal chemical industry topology knowledge graph is then time-series encoded according to the time window and process unit to obtain dynamic encoding.

[0067] A coal chemical industry standard metadata storage model is constructed based on the dynamic encoding. The data to be stored is input into the coal chemical industry standard metadata storage model, and the storage result is output.

[0068] In this embodiment, the method for extracting information data from the associated data using multimodal semantic information includes:

[0069] Extract general information: Extract process-related information from associated data to obtain equipment performance characteristics, material conversion patterns, and energy consumption coupling relationships; perform semantic annotation on document data to extract explicit and implicit knowledge to obtain key points of process procedures, analysis report conclusions, and design specification requirements; perform visual semantic extraction on multimedia data to transform it into describable information to obtain equipment status indicators, abnormal operating condition indicators, and environmental feature extraction.

[0070] Incremental information extraction: Extract equipment tag number, medium name, and operating pressure triplets from flow charts; extract sample number, test items, values, and compliance standards from analysis reports; extract step numbers, operating actions, and process parameters from operating procedures; directly extract incremental information from document data; for multimedia data: extract location, corrosion area, and severity from reactor lining corrosion photos in equipment status diagrams; perform OCR recognition on instrument reading images to obtain pressure gauge and thermometer readings and associate them with corresponding measurement point codes; extract keyframes from tank area monitoring videos to obtain liquid level, time, and abnormal status; identify maintenance operation videos to obtain tool type, operating steps, and personnel qualifications;

[0071] Using equipment, materials, processes, and environment from general information as core entities, establish multi-dimensional relationships; perform semantic association based on incremental information to obtain image-real-time data association and video-fault tracing association;

[0072] When multi-source data encounters information contradictions, a three-level verification method is used for verification: priority determination, time correlation, and credibility weighting. Priority determination is: online analytical instrument data > laboratory test reports > manual records. Time correlation prioritizes multimodal data within the same time window. Credibility weighting assigns credibility coefficients to data from different sources, and the weighted values ​​are used for fusion when conflicts occur.

[0073] The multi-dimensional correlation, image-real-time data correlation, and video-fault tracing correlation are output as information data.

[0074] In the actual evaluation, the equipment performance characteristics are the three-dimensional correlation curve of gasifier temperature-pressure-load, based on the aggregation of historical data of the past 3 months from the TDengine time series database; the material conversion law is the influence of coal type change on synthesis reaction, which is correlated with coal quality analysis report and online chromatographic data; the energy consumption coupling relationship is the oxygen production-electricity consumption curve of air separation unit, combined with power monitoring system and production scheduling records.

[0075] The key points of the process specifications are: extracting the binding clause from the "Gasifier Operation Specifications" that the start-up heating rate must not exceed 50℃ / h, and referring to the material parameters in the equipment maintenance manual; the analysis report conclusion is to extract the key conclusion from the catalyst deactivation analysis report that the deactivation rate of iron-based catalysts is accelerated by 30% when the sulfur content is >0.1ppm, and adding the page numbers of the experimental data source; the design specification requirements are to extract the design parameter of reactor lining corrosion allowance ≥3mm from the "Coal-to-Oil Plant Design Manual", and verify the compliance of the data on the equipment nameplate.

[0076] The status identifier is marked as the corrosion area in the reactor lining photo, with the shooting date and location coordinates; the abnormal operating condition identifier is the key frame description of the abnormal liquid level fluctuation event in the tank area monitoring video, associated with the DCS alarm record; the environmental feature extraction is the visual features of pipe rack insulation layer damage and valve leakage frost in the plant area drone inspection image, and locates the specific equipment tag number.

[0077] Equipment-process correlation: Gasifier → Corresponding process parameters → Operating constraints; Material-quality correlation: Raw coal → Key components → Impact on product quality; Anomaly-handling correlation: Synthesis tower pressure drop alarm → Possible causes → Emergency handling steps;

[0078] Catalyst test report: The original content fragment states that the activity of this batch of CAT-200 catalyst is ≥92%, and the recommended replacement cycle is 8000h; the extracted incremental information is related to equipment R-101 (gasifier); the activity is 92%, and the replacement cycle is 8000h;

[0079] Environmental impact assessment approval document: The original content fragment is SO2 emission limit: 100mg / m³ (monitoring point: exhaust stack DA001); the extracted incremental information is environmental protection indicator SO2=100mg / m³, monitoring point DA001 (related to GB 16171-2012).

[0080] The credibility weighting is calculated based on a comprehensive assessment of the historical accuracy of the data source, the device calibration status, and the sampling frequency.

[0081] General information: Extract the temperature-pressure coupling relationship of the gasifier from the associated data, extract the key point of pressurization rate ≤0.1MPa / min from the "Synthesis Tower Operation Procedure", and identify abnormal operating conditions such as valve leakage from the equipment inspection images;

[0082] Incremental information: Extract the equipment tag number R-101, medium name crude syngas, and operating pressure 3.5MPa ternary group from the process flow diagram; identify the pressure gauge image through OCR to obtain a reading of 2.8MPa and associate it with the measuring point code PT-001; extract information from the maintenance video that the tool type is a wrench and the operation step is tightening bolts;

[0083] Data verification: When online instrument data conflicts with laboratory test reports, the online instrument data shall prevail; multimodal data within the same time window shall be correlated, and the online instrument data shall be assigned a confidence coefficient of 0.9, the laboratory report 0.7, and the manual record 0.5. In case of conflict, the data shall be merged according to the weighted values.

[0084] In this embodiment, the method for obtaining structured data by structuring the information data and the job data includes:

[0085] For information data: Through three steps of extraction, cleaning, and association, key information is extracted, missing values ​​are handled, and outliers are corrected. Based on the information data, a chain traceability is achieved through foreign key associations to realize raw coal batches → quality inspection reports → supplier information. Relationship modeling is performed based on the chain traceability to obtain structured information data.

[0086] A three-level coding rule is used to assign a unique identifier to each monitoring point in the operation data; the five elements of the operation data, namely equipment number, data category, timestamp, value, and quality mark, are recorded according to the monitoring point to obtain structured operation data; the database is partitioned by process unit and day to obtain hot data and cold data; the hot data area stores the current and the most recent day's second / minute level data, the warm data area stores downsampled data from 1 to 30 days, and the cold data area stores archived data from more than 30 days.

[0087] Attach association pointers to unstructured files, obtain the relationship between equipment, materials, and parameters based on structured information data and structured operational data, construct a structured knowledge graph based on the relationship between equipment, materials, and parameters, record the entire data lifecycle, and obtain cross-modal data associations;

[0088] Output the structured knowledge graph as structured data;

[0089] In actual assessment, information and data are used to extract key information such as catalyst activity (≥92%) and replacement cycle (8000h) from the analysis report, fill in missing coal quality analysis data using interpolation, and correct abnormal pressure values ​​that exceed the value range; and achieve chain traceability of raw coal batch B20240501 → quality inspection report QR-20240501 → supplier S-003 through foreign key association.

[0090] Operational data: A three-level coding rule of process unit code + equipment code + measuring point code is adopted to assign a unique identifier to PT-001 (gasifier pressure measuring point); five elements are recorded according to the monitoring point: equipment number, data category, timestamp, value, and quality mark; the data is divided into process unit + day level partitions, with the hot data area storing the second-level data of the current day, the warm data area storing the minute-level average of 1-30 days, and the cold data area storing archived data of more than 30 days.

[0091] Cross-modal association: Add association pointers to the corrosion photos of the reactor lining, link them to equipment R-101 and corresponding process parameters, build a structured knowledge graph containing the relationship between equipment, materials and parameters, and record the entire life cycle of data from acquisition to archiving.

[0092] In this embodiment, the method for constructing a process mapping table based on business terms and process parameters includes:

[0093] Based on business terminology, key parameters are extracted from the process parameters. These key parameters include process unit, process parameter, number type, unit, value range, associated equipment ID, and data source system. A process mapping table is then constructed based on the business terminology and key parameters.

[0094] Establish a dynamic update mechanism. After accepting an update request, modify the process mapping table in the data management platform and update the process unit, process parameters, number type, unit, value range, associated equipment ID, and data source system synchronously. After the update, verify the consistency of the updated data. The process mapping table adopts the naming rule of basic version and change number. Each update records the person making the change, the time of the change, and the content of the change.

[0095] When process parameters change, business terms are added, or standards and specifications are updated, a dynamic update mechanism is triggered to update the process mapping table and output the updated process mapping table.

[0096] In actual evaluation, key parameters are extracted as follows: Based on the coal-to-ethylene industry terminology database, key parameters such as process unit (gasification unit), process parameters (gasification temperature), number type (floating point), unit (°C), value range (800-1400), associated equipment ID (R-101), and data source system (DCS) are extracted.

[0097] Dynamic update: When the oxygen-to-coal ratio of the process parameter is adjusted to 0.9-1.3, the update mechanism is triggered. The process mapping table is modified in the data management platform, the person making the change and the time of the change are recorded, the associated equipment ID and the data source system are updated synchronously, and the consistency of the updated data with other parameters is verified.

[0098] In this embodiment, the method for constructing a knowledge graph of coal chemical industry topology includes:

[0099] With process flow, material flow, and energy flow as the core, the process units, equipment, materials, and parameter elements of the entire coal chemical process are abstracted into nodes through entity and relation modeling, and directed edges are used to represent the logical relationships between elements.

[0100] Entity types are categorized based on the characteristics of structured data and process logic. Each entity type contains standardized attributes, including process units, equipment, materials, process parameters, operating conditions, and analysis indicators.

[0101] Based on the definition of chemical process logic, the semantic relationship between source entities and target entities is clarified. Combining the characteristics of coal chemical process, core association types are extracted, including inclusion, feeding, generation, control, consumption, satisfaction, association, and flow direction. The direction of inclusion is from process unit to equipment; the direction of feeding is from material to equipment; the direction of generation is from equipment to material; the direction of control is from process parameters to equipment; the direction of consumption is from equipment to material; the direction of satisfaction is from equipment to operating conditions; the direction of association is from analytical indicators to material; and the direction of flow direction is from material to process unit.

[0102] The knowledge graph core framework is constructed based on entity type and core association type. The topology structure is decomposed according to the five major process units of raw material pretreatment, gasification, purification, synthesis and separation. The connection logic of elements within and between units is clarified to obtain the core process topology structure of coal chemical industry.

[0103] Construct the topology of the raw material pretreatment unit, and obtain the core entities and relationships based on the entity types of the raw material pretreatment unit. The entity types include process unit, equipment, material and process parameters. The core entity of the process unit is the raw material pretreatment unit, the core entities of the equipment are crushing equipment and dryer, the core entities of the material are raw coal and crushed coal, and the core entities of the process parameters are crushing particle size and drying temperature.

[0104] Construct the topology of the gasification unit, and obtain the core entities and relationships through the entity types of the gasification unit. The entity types include process unit, equipment, material and operation unit. The core entity of the process unit is the gasification unit, the core entities of the equipment are the gasifier and burner, the core entity of the material is the crude syngas, and the core entities of the process parameters are the gasifier temperature and oxygen-coal ratio.

[0105] Construct the purification unit topology, and obtain the core entities and relationships through the entity types of the purification units. The entity types include process units, equipment, materials and operation units. The core entity of the process unit is the purification unit, the core entity of the equipment is the shift reactor, the core entity of the material is the purified gas, and the core entity of the process parameters is the catalyst activity conditions.

[0106] Construct the topology of synthesis and separation units, and obtain the core entities and relationships based on the entity types of synthesis and separation units. Entity types include process units, equipment, materials and analytical indicators. The core entities of process units are synthesis units and separation units, the core entities of equipment are synthesis tower and ethylene distillation tower, the core entity of materials is polymer-grade ethylene, and the core entity of analytical indicators is ethylene purity.

[0107] Based on the core framework of the knowledge graph and the core process topology of coal chemical industry, a knowledge graph of coal chemical industry topology structure is output.

[0108] In actual evaluation, entity and relationship modeling is used: process units, equipment, materials, etc. are abstracted as nodes, and directed edges are used to represent the relationships; for example: gasification unit → contains → R-101 gasifier, raw coal → feed → R-101 gasifier, R-101 gasifier → generate → crude syngas;

[0109] Unit topology construction: Raw material pretreatment unit: The core entities include crushing equipment, dryer, raw coal, and crushed coal. The relationship is raw coal → feed → crushing equipment → production → crushed coal → feed → dryer; Synthesis unit: The core entities are synthesis tower and polymerization-grade ethylene. The process parameters are reaction temperature (200-250℃). The relationship is purified gas → feed → synthesis tower → production → polymerization-grade ethylene.

[0110] Graph Output: Integrating the topologies of the five major process units to form a complete knowledge graph of coal chemical industry topology, covering 1200+ nodes and 3000+ relationships.

[0111] In this embodiment, the method for obtaining dynamic coding by performing time-series coding on the coal chemical industry topology knowledge graph includes:

[0112] Historical data is acquired to construct a static coding baseline. A hierarchical dynamic coding system is then employed, fusing process unit attributes, time window features, and data popularity weights. Priority is dynamically adjusted using a popularity evaluation function, the expression of which is:

[0113] ;

[0114] in For heat evaluation function, This is the process correlation coefficient. This represents the current data query frequency. This represents the system's maximum query frequency. The standard deviation of the data. The maximum allowable standard deviation of process variation. The weighting coefficient for access frequency. This is a weighting coefficient for process correlation. The weighting coefficient for volatility sensitivity;

[0115] Historical data is used to train the process correlation coefficient and optimize parameters. Given a priority dynamic adjustment rule, when the heat evaluation value is greater than or equal to 0.802, the encoding priority is the highest, and an independent memory cache is allocated; when the heat evaluation value is greater than or equal to 0.51 and less than 0.802, the priority is medium, and the data is input into the hot data pool and updated hourly; when the heat evaluation value is less than 0.51, the priority is low, and the data is input into the cold data archive.

[0116] A time window sliding mechanism is set up: a short window is used to store high-frequency data at the second level, and the complete timestamp is retained during encoding; a medium window is used to downsample to the minute-level mean, and similar detection points are merged during encoding; a long window is used to compress hour-level feature values, and only key indicators are retained during encoding.

[0117] Deploy time window sliding triggers based on time window sliding mechanism, and perform cross-unit data association through process knowledge graph edge attribute encoding. The edge attribute fields are composed of source unit code, target unit code, association type code and time delay in sequence, and the hierarchical dynamic encoding is output as dynamic encoding.

[0118] In actual evaluation, multiple anomaly detection thresholds are obtained: to identify isolated entities, anomaly thresholds are set based on the statistical distribution of historical data, and automatic review is triggered when the anomaly score is >μ+2σ (where μ is the mean of historical anomaly scores and σ is the standard deviation); normal connection patterns are learned through GNN, and anomalies are judged when the reconstruction error is >Q3+1.5×IQR (Q3 is the third quartile and IQR is the interquartile range).

[0119] The hierarchical dynamic coding adopts an 18-bit fixed-length code, which consists of: 6 bits for process unit identifier, 4 bits for time window code, 3 bits for data type code, 3 bits for heat priority, and 2 bits for check bit;

[0120] The process correlation coefficient C is trained using historical data, and the parameters are optimized using gradient descent.

[0121] The process correlation coefficient is obtained through historical optimization and theoretical training, with a short-term window of 1 day, a medium-term window of 1 day, and a long-term window of 90 days.

[0122] The edge attribute fields are: [Source cell code (18 bits)][Target cell code (18 bits)][Association type code (2 bits)][Time delay (4 bits)];

[0123] Static coding baseline: built based on the project's historical data over the past year, with the process correlation coefficient set at 0.7 (optimized through training with historical data).

[0124] Popularity Assessment: The current query frequency of a certain monitoring point is 50 times / day, the maximum query frequency of the system is 100 times / day, the data standard deviation is 5, the maximum allowable standard deviation of process fluctuation is 10, the weight coefficients W1=0.3, W2=0.4, W3=0.3, the calculated popularity assessment value H=0.3×(50 / 100)+0.4×0.7+0.3×(5 / 10)=0.63, which is determined to be of medium priority and stored in the hot data pool and updated hourly;

[0125] Time windows: Short-term window (1 day) stores second-level data and retains complete timestamps; medium-term window (30 days) stores minute-level averages; long-term window (90 days) stores hour-level feature values, retaining only key indicators such as gasification temperature and ethylene purity.

[0126] Perception Layer: Deploys a BERT-based industrial fine-tuning model to identify equipment name R-101, parameter vaporization temperature, and unit ℃ triples, automatically constructing an asset hierarchy tree; Storage Layer: Uses Neo4j graph database to store the knowledge graph, TDengine time-series library to store dynamic metadata, and MinIO object storage to store associated unstructured files, achieving cross-database association through unified encoding; Decision Layer: Runs anomaly detection algorithms, determining an anomaly when the reconstruction error of a certain measurement point data is >0.3, automatically triggering review; metadata is synchronously updated to the associated system every 15 minutes; Application Layer: Based on similarity search algorithms, outputs services such as equipment fault tracing and process parameter optimization, feeding back feedback data to the perception layer to form a closed loop.

[0127] In this embodiment, the method for constructing a coal chemical industry standard metadata storage model based on the dynamic encoding includes:

[0128] The coal chemical industry standard metadata storage model comprises four modules: perception, storage, decision-making, and application. It achieves dynamic metadata growth through an algorithm chain. The perception layer deploys entity recognition and natural language processing algorithms to extract entities, attributes, and relationships from multi-source heterogeneous data sources, forming initial triples. The storage layer uses the Neo4j graph database to store the knowledge graph, a time-series database to store dynamic metadata, and object storage to associate unstructured files, achieving cross-database association through unified encoding. The decision-making layer runs anomaly detection and clustering / classification algorithms to verify metadata quality in real time and dynamically optimize the knowledge graph structure. The application layer, based on similarity search and link prediction algorithms, outputs business services and feeds back feedback data to the perception layer, forming a self-iteratory closed loop.

[0129] Develop dynamic coding rules and attach lifecycle status codes and association strength weights to each metadata element. The dynamic coding rules include four layers: metadata collection and construction, metadata storage and indexing, metadata quality and intelligent management, and metadata application and value mining.

[0130] Metadata collection and construction: A hybrid strategy of rule-based and deep learning is adopted to address the multi-source nature of equipment names. Rule cleaning is performed based on an industry terminology database, string similarity is calculated and embedded into a knowledge graph. When a new entity is added to the database, its contextual relationship is learned through node embedding vectors. After comparison with existing entity vectors, it is automatically classified. Industrial fine-tuning is performed based on the BERT-base model to identify equipment name, parameter, and unit triplet. Few-shot learning is used to process scarce labeled data. Relationships are identified through trigger words and dependency syntax, and an asset hierarchy tree is automatically constructed.

[0131] Metadata Storage and Indexing: Neo4j graph database is used as the core storage engine. Graph algorithms are used to achieve efficient organization and dynamic querying of metadata relationships. Key encoding strategies include dynamic graph structure optimization, community detection, and dynamic clustering. Time-series metadata is stored in partitions by device ID and time window. Through APOC plugin linkage, millisecond-level association queries of real-time data, static attributes, and relationship networks are achieved. Graph structure dynamic optimization involves hierarchical storage of core entities, related entities, and attribute entities. Core entities use memory-first indexing, and weight values ​​are added to directed edges to support Dijkstra's algorithm for finding critical dependency paths. Community detection and dynamic clustering identify tight clusters in the metadata graph through community detection algorithms and automatically generate process unit knowledge packages.

[0132] Metadata Quality and Intelligent Management: A three-layer defense is deployed to address common metadata quality issues, identifying isolated entities and automatically triggering review when the anomaly score > 0.85; GNN-based reconstruction of normal connection patterns with an error > 0.3 is used to classify anomalies; historical metadata change records are compared, and if the timestamp deviation from the process change order is > 24 hours, it is automatically marked as pending synchronization; for new metadata, semi-supervised learning is used for automatic classification, dividing new measurement points into multiple initial clusters based on collection frequency, engineering unit, and region, and training a model based on existing labels; when metadata is updated, it is automatically pushed to multiple related systems via graph query statements, with a synchronization delay ≤ 15 minutes;

[0133] The metadata application and value mining process consists of four layers: comparing equipment type, process location, and key parameters for attribute matching; using graph neural networks to aggregate neighbor information and output a comprehensive similarity score; learning the embedded representation of known metadata relationships through graph neural networks; and using node similarity and structural features, where metadata relationships include causal relationships and process optimization.

[0134] In the actual evaluation, the object was MinIO, and the initial cluster size was 8.

[0135] Figure 2 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 2 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0136] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 2 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0137] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0138] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a standard metadata storage device at the logical level. The processor executes the program stored in the memory and specifically performs any of the aforementioned standard metadata storage methods.

[0139] The above is as stated in this application. Figure 1The standard metadata storage method disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0140] The electronic device can also perform Figure 1 A standard metadata storage method is proposed and implemented. Figure 1 The functions of the embodiments shown are not described in detail here.

[0141] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, perform any of the aforementioned standard metadata storage methods.

[0142] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0143] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0144] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0145] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0146] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0147] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0148] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0149] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0150] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0151] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A standard metadata storage method, characterized in that, Includes the following steps: Collect pre-defined operational data and related data for coal chemical engineering, and preprocess the operational data and related data; the operational data includes process parameters, process condition data, equipment static attributes, and material and energy data; the related data includes document data and multimedia data; the document data includes design documents, analysis reports, and compliance documents; the multimedia data includes equipment images and video streams; Multimodal semantic information extraction is performed on the associated data to obtain information data; the information data and the operation data are structured to obtain structured data; and a process mapping table is constructed based on business terms and process parameters; the information data includes incremental information and general information. Based on chemical industry association types, a coal chemical industry topology knowledge graph is constructed using the structural data and the process mapping table. Dynamic encoding is obtained by temporal encoding of the coal chemical industry topology knowledge graph according to time windows and process units; including: Historical data is acquired to construct a static coding baseline. A hierarchical dynamic coding system is then employed, fusing process unit attributes, time window features, and data popularity weights. Priority is dynamically adjusted using a popularity evaluation function, the expression of which is: in For heat evaluation function, This is the process correlation coefficient. This represents the current data query frequency. This represents the system's maximum query frequency. The standard deviation of the data. The maximum allowable standard deviation of process variation. The weighting coefficient for access frequency. This is a weighting coefficient for process correlation. The weighting coefficient for volatility sensitivity; Historical data is used to train the process correlation coefficient and optimize parameters. Given a priority dynamic adjustment rule, when the heat evaluation value is greater than or equal to 0.802, the encoding priority is the highest, and an independent memory cache is allocated; when the heat evaluation value is greater than or equal to 0.51 and less than 0.802, the priority is medium, and the data is input into the hot data pool and updated hourly; when the heat evaluation value is less than 0.51, the priority is low, and the data is input into the cold data archive. A time window sliding mechanism is set up: a short window is used to store high-frequency data at the second level, and the complete timestamp is retained during encoding; a medium window is used to downsample to the minute-level mean, and similar detection points are merged during encoding; a long window is used to compress hour-level feature values, and only key indicators are retained during encoding. Deploy time window sliding triggers based on time window sliding mechanism, and perform cross-unit data association through process knowledge graph edge attribute encoding. The edge attribute fields are composed of source unit code, target unit code, association type code and time delay in sequence, and the hierarchical dynamic encoding is output as dynamic encoding. A coal chemical industry standard metadata storage model is constructed based on the dynamic encoding. The data to be stored is input into the coal chemical industry standard metadata storage model, and the storage result is output.

2. The standard metadata storage method according to claim 1, characterized in that, A method for extracting information data from the associated data using multimodal semantic information extraction includes: Extract general information: Extract process-related information from associated data to obtain equipment performance characteristics, material conversion patterns, and energy consumption coupling relationships; perform semantic annotation on document data to extract explicit and implicit knowledge to obtain key points of process procedures, analysis report conclusions, and design specification requirements; perform visual semantic extraction on multimedia data to transform it into describable information to obtain equipment status indicators, abnormal operating condition indicators, and environmental feature extraction. Incremental information extraction: Extract equipment tag number, medium name, and operating pressure triplets from flow charts; extract sample number, test items, values, and compliance standards from analysis reports; extract step numbers, operating actions, and process parameters from operating procedures; directly extract incremental information from document data; for multimedia data: extract location, corrosion area, and severity from reactor lining corrosion photos in equipment status diagrams; perform OCR recognition on instrument reading images to obtain pressure gauge and thermometer readings and associate them with corresponding measurement point codes; extract keyframes from tank area monitoring videos to obtain liquid level, time, and abnormal status; identify maintenance operation videos to obtain tool type, operating steps, and personnel qualifications; Using equipment, materials, processes, and environment from general information as core entities, establish multi-dimensional relationships; perform semantic association based on incremental information to obtain image-real-time data association and video-fault tracing association; When multi-source data encounters information contradictions, a three-level verification method is used for verification: priority determination, time correlation, and credibility weighting. In priority determination, online analysis instrument data > laboratory test reports > manual records. Time correlation prioritizes multimodal data within the same time window. Credibility weighting assigns credibility coefficients to data from different sources, and the weighted values ​​are used for fusion when conflicts occur. The multi-dimensional correlation, image-real-time data correlation, and video-fault tracing correlation are output as information data.

3. The standard metadata storage method according to claim 1, characterized in that, A method for obtaining structured data by structuring the information data and the job data includes: For information data: Through three steps of extraction, cleaning, and association, key information is extracted, missing values ​​are handled, and outliers are corrected. Based on the information data, a chain traceability is achieved through foreign key associations to realize raw coal batches → quality inspection reports → supplier information. Relationship modeling is performed based on the chain traceability to obtain structured information data. A three-level coding rule is used to assign a unique identifier to each monitoring point in the operation data; the five elements of the operation data, namely equipment number, data category, timestamp, value, and quality mark, are recorded according to the monitoring point to obtain structured operation data; the database is partitioned by process unit and day to obtain hot data and cold data; the hot data area stores the current and the most recent day's second / minute level data, the warm data area stores downsampled data from 1 to 30 days, and the cold data area stores archived data from more than 30 days. Attach association pointers to unstructured files, obtain the relationship between equipment, materials, and parameters based on structured information data and structured operational data, construct a structured knowledge graph based on the relationship between equipment, materials, and parameters, record the entire data lifecycle, and obtain cross-modal data associations; Output the structured knowledge graph as structured data.

4. The standard metadata storage method according to claim 1, characterized in that, A method for constructing a process mapping table based on business terms and process parameters includes: Based on business terminology, key parameters are extracted from the process parameters. These key parameters include process unit, process parameter, number type, unit, value range, associated equipment ID, and data source system. A process mapping table is then constructed based on the business terminology and key parameters. Establish a dynamic update mechanism. After accepting an update request, modify the process mapping table in the data management platform and update the process unit, process parameters, number type, unit, value range, associated equipment ID, and data source system synchronously. After the update, verify the consistency of the updated data. The process mapping table adopts the naming rule of basic version and change number. Each update records the person making the change, the time of the change, and the content of the change. When process parameters change, business terms are added, or standards and specifications are updated, a dynamic update mechanism is triggered to update the process mapping table and output the updated process mapping table.

5. The standard metadata storage method according to claim 1, characterized in that, Methods for constructing knowledge graphs of coal chemical industry topology include: With process flow, material flow, and energy flow as the core, the process units, equipment, materials, and parameter elements of the entire coal chemical process are abstracted into nodes through entity and relation modeling, and directed edges are used to represent the logical relationships between elements. Entity types are categorized based on the characteristics of structured data and process logic. Each entity type contains standardized attributes, including process units, equipment, materials, process parameters, operating conditions, and analysis indicators. Based on the definition of chemical process logic, the semantic relationship between source entities and target entities is clarified. Combining the characteristics of coal chemical process, core association types are extracted, including inclusion, feeding, generation, control, consumption, satisfaction, association, and flow direction. The direction of inclusion is from process unit to equipment; the direction of feeding is from material to equipment; the direction of generation is from equipment to material; the direction of control is from process parameters to equipment; the direction of consumption is from equipment to material; the direction of satisfaction is from equipment to operating conditions; the direction of association is from analytical indicators to material; and the direction of flow direction is from material to process unit. The knowledge graph core framework is constructed based on entity type and core association type. The topology structure is decomposed according to the five major process units of raw material pretreatment, gasification, purification, synthesis and separation. The connection logic of elements within and between units is clarified to obtain the core process topology structure of coal chemical industry. Construct the topology of the raw material pretreatment unit, and obtain the core entities and relationships based on the entity types of the raw material pretreatment unit. The entity types include process unit, equipment, material and process parameters. The core entity of the process unit is the raw material pretreatment unit, the core entities of the equipment are crushing equipment and dryer, the core entities of the material are raw coal and crushed coal, and the core entities of the process parameters are crushing particle size and drying temperature. Construct the topology of the gasification unit, and obtain the core entities and relationships through the entity types of the gasification unit. The entity types include process unit, equipment, material and operation unit. The core entity of the process unit is the gasification unit, the core entities of the equipment are the gasifier and burner, the core entity of the material is the crude syngas, and the core entities of the process parameters are the gasifier temperature and oxygen-coal ratio. Construct the purification unit topology, and obtain the core entities and relationships through the entity types of the purification units. The entity types include process units, equipment, materials and operation units. The core entity of the process unit is the purification unit, the core entity of the equipment is the shift reactor, the core entity of the material is the purified gas, and the core entity of the process parameters is the catalyst activity conditions. Construct the topology of synthesis and separation units, and obtain the core entities and relationships based on the entity types of synthesis and separation units. Entity types include process units, equipment, materials and analytical indicators. The core entities of process units are synthesis units and separation units, the core entities of equipment are synthesis tower and ethylene distillation tower, the core entity of materials is polymer-grade ethylene, and the core entity of analytical indicators is ethylene purity. Based on the core framework of the knowledge graph and the core process topology of coal chemical industry, a knowledge graph of coal chemical industry topology structure is output.

6. The standard metadata storage method according to claim 1, characterized in that, The method for constructing a coal chemical standard metadata storage model based on the dynamic encoding includes: The coal chemical industry standard metadata storage model comprises four modules: perception, storage, decision-making, and application. It achieves dynamic metadata growth through an algorithm chain. The perception layer deploys entity recognition and natural language processing algorithms to extract entities, attributes, and relationships from multi-source heterogeneous data sources, forming initial triples. The storage layer uses the Neo4j graph database to store the knowledge graph, a time-series database to store dynamic metadata, and object storage to associate unstructured files, achieving cross-database association through unified encoding. The decision-making layer runs anomaly detection and clustering / classification algorithms to verify metadata quality in real time and dynamically optimize the knowledge graph structure. The application layer, based on similarity search and link prediction algorithms, outputs business services and feeds back feedback data to the perception layer, forming a self-iteratory closed loop. Develop dynamic coding rules and attach lifecycle status codes and association strength weights to each metadata element. The dynamic coding rules include four layers: metadata collection and construction, metadata storage and indexing, metadata quality and intelligent management, and metadata application and value mining. Metadata collection and construction: A hybrid strategy of rule-based and deep learning is adopted to address the multi-source nature of equipment names. Rule cleaning is performed based on an industry terminology database, string similarity is calculated and embedded into a knowledge graph. When a new entity is added to the database, its contextual relationship is learned through node embedding vectors. After comparison with existing entity vectors, it is automatically classified. Industrial fine-tuning is performed based on the BERT-base model to identify equipment name, parameter, and unit triplet. Few-shot learning is used to process scarce labeled data. Relationships are identified through trigger words and dependency syntax, and an asset hierarchy tree is automatically constructed. Metadata Storage and Indexing: Neo4j graph database is used as the core storage engine. Graph algorithms are used to achieve efficient organization and dynamic querying of metadata relationships. Key encoding strategies include dynamic graph structure optimization, community detection, and dynamic clustering. Time-series metadata is stored in partitions by device ID and time window. Through APOC plugin linkage, millisecond-level association queries of real-time data, static attributes, and relationship networks are achieved. Graph structure dynamic optimization involves hierarchical storage of core entities, related entities, and attribute entities. Core entities use memory-first indexing, and weight values ​​are added to directed edges to support Dijkstra's algorithm for finding critical dependency paths. Community detection and dynamic clustering identify tight clusters in the metadata graph through community detection algorithms and automatically generate process unit knowledge packages. Metadata Quality and Intelligent Management: A three-layer defense is deployed to address common metadata quality issues, identifying isolated entities and automatically triggering review when the anomaly score > 0.85; GNN-based reconstruction of normal connection patterns with an error > 0.3 is used to classify anomalies; historical metadata change records are compared, and if the timestamp deviation from the process change order is > 24 hours, it is automatically marked as pending synchronization; for new metadata, semi-supervised learning is used for automatic classification, dividing new measurement points into multiple initial clusters based on collection frequency, engineering unit, and region, and training a model based on existing labels; when metadata is updated, it is automatically pushed to multiple related systems via graph query statements, with a synchronization delay ≤ 15 minutes; Metadata application and value mining: Attribute matching is performed by comparing equipment type, process location, and key parameters. Graph neural network aggregates neighbor information to output a comprehensive similarity score. Embedded representations of known metadata relationships are learned through graph neural network, based on node similarity and structural features. Metadata relationships include causal relationships and process optimization.

7. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method according to any one of claims 1 to 6.

8. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-source data processing system for geographic information big data

    CN120353874A

  • Power distribution cabinet maintenance system based on artificial intelligence

    CN120855652A