Environmental corrosion data storage and management system

By integrating multi-source heterogeneous data and adopting advanced data processing and management methods, the problem of data silos in the environmental corrosion monitoring data management system has been solved, data fusion and unification have been achieved, the depth of data cleaning and the quality of reports have been improved, and an independent and controllable domestic platform has been built.

CN121657950APending Publication Date: 2026-03-13CHINESE PEOPLES LIBERATION ARMY NAVAL SERVICE ACAD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing environmental corrosion monitoring data management systems, massive monitoring data resources are scattered, data sharing and business collaboration lack overall planning, and data integration and unification have not been achieved.

Method used

The system integrates multi-source heterogeneous data using a data acquisition and reception module, combines Flink and Hive's integrated stream and batch dynamic processing methods, Grubbs test and machine learning unsupervised models for data cleaning and auditing, constructs a corrosion monitoring knowledge graph, and generates intelligent thematic reports through graph neural networks to achieve multi-dimensional data asset management.

Benefits of technology

It effectively integrates data resources from multiple fields, solves the problem of data silos, achieves data fusion and unification, enhances the depth of data cleaning and the decision-making value of reports, provides a scientific basis for data asset operation, and builds an independent and controllable domestic platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121657950A_ABST
    Figure CN121657950A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage management, in particular to an environment corrosion data storage and management system which comprises a data acquisition and connection module, a data governance module and a data management module. The data acquisition and connection module is used for integrating multi-source heterogeneous data and realizing online connection and offline filling of the data; and the data governance module comprises a cleaning and auditing unit which performs data cleaning and auditing by adopting a flow batch integrated dynamic processing method based on Flink and Hive in combination with a Grubbs inspection and machine learning unsupervised model. According to the method, multi-source heterogeneous corrosion data is integrated, Flink and Hive flow batch integrated processing is carried out, and Grubbs inspection and a machine learning unsupervised model are combined to clean auditing data; the core level is replaced in a localization manner; a corrosion monitoring knowledge graph is constructed, an intelligent report is generated by means of a G2T model of GNN, the problem of data islands is solved, and data fusion and unification are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data storage management technology, specifically to an environmental corrosion data storage and management system. Background Technology

[0002] In high-temperature, high-humidity, and high-salt environments, the efficient collection and scientific management of environmental corrosion monitoring data covering multiple fields such as meteorology, ecology, geology, hydrodynamics, and corrosion can effectively improve the service safety and life-cycle protection of concrete structures, metal structures, and engineering equipment. It is noteworthy that the rapid development of domestically produced platforms has brought breakthrough opportunities for data storage and management, effectively enabling the domestic independent control of environmental corrosion data storage and management systems.

[0003] However, in practical applications, existing environmental corrosion monitoring data management systems are mostly built independently on the Windows platform, resulting in a large amount of scattered monitoring data resources, serious "data silos," and a lack of overall planning for data sharing and business collaboration, thus failing to achieve data integration and unification. Summary of the Invention

[0004] The purpose of this invention is to provide an environmental corrosion data storage and management system to solve the problems of scattered massive monitoring data resources, lack of overall planning in data sharing and business collaboration, and failure to achieve data integration and unification.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an environmental corrosion data storage and management system, comprising: It includes a data acquisition and reception module, a data governance module, and a data management module; The data acquisition and reception module is used to integrate multi-source heterogeneous data from meteorological, ecological environment, geological, hydrodynamic and facility corrosion fields to realize online data reception and offline data entry. The data governance module includes a cleaning and auditing unit that uses a dynamic processing method based on Flink and Hive, combined with Grubbs test and machine learning unsupervised model for data cleaning and auditing, and a metadata management unit that constructs a corrosion monitoring knowledge graph, and constructs the corrosion monitoring knowledge graph through metadata management. The data management module generates intelligent thematic reports based on the graph-to-text (G2T) model of graph neural networks (GNN), and uses a multi-dimensional data asset health assessment model (including basic quality, structural correlation, and value service dimensions) for data asset management.

[0006] Preferably, the online data transfer specifically involves establishing a database connection through the DataX component to automatically acquire data from multiple heterogeneous data sources in real time or periodically, ensuring the timeliness and accuracy of the data. The online data transfer specifically includes: Data source management unit: It adopts a web-based system based on DataX components that is compatible with multiple database connections. It is used to connect to multiple databases with multiple data sources simultaneously, and realizes online data collection from each database through the design of data call interfaces. Task Management Unit: Used for project management, data collection templates, task creation, visual data collection task orchestration, setting task parameters, and timed data collection from database tables.

[0007] Preferably, the data governance module further includes a data quality standard unit and a data quality unit: The data quality standard unit defines scoring rules for completeness, validity, and uniqueness; The data quality unit performs data quality monitoring, anomaly alarms, and quality analysis report generation based on the data quality standard unit.

[0008] Preferably, the metadata management unit collects table structure and view definition metadata through DataX to complete the unified storage of metadata; By parsing metadata relationships using Apache Atlas, a data lineage graph can be constructed. A knowledge graph for corrosion monitoring is constructed using a graph data model. The nodes of the knowledge graph include monitoring facilities, environmental parameters, corrosion events, and quality reports, while the edges include physical associations, logical dependencies, and corrosion driving factors.

[0009] Preferably, the data governance module's cleaning and auditing unit has a rule base that stores cleaning rules for null value filling, data standardization, unit conversion, and outlier smoothing, as well as auditing rules for numerical rules, formula rules, and format rules; the machine learning unsupervised model is an autoencoder or IsolationForest, used to identify time-series data pattern deviations and sensor malfunctions.

[0010] Preferably, the multi-dimensional data asset health assessment model includes: a basic quality dimension including the proportion of outliers; a structural and relational dimension including the integrity of the lineage and the cross-topic relevance index; and a value and service dimension including real-time stream acquisition latency and API call success rate.

[0011] Preferably, in the process of generating the intelligent special report, the graph neural network model uses a graph attention network (GAT) or a graph Transformer as an encoder to extract knowledge graph association features, and generates conclusions and suggestions containing logical reasoning and an assessment of the ecological environment status through a Transformer or RNN / LSTM decoder.

[0012] Preferably, the data management module includes: Data editing unit: used to support the editing of time-series and attribute data; Data query unit: used to support multiple query methods, including map query; Data Asset Management Unit: Used for monitoring and statistical analysis of global data assets; Special topic data statistics unit: used to support multi-dimensional data statistics; Special Topic Report Generation Unit: Used to generate analysis reports based on the reporting standards for each topic; Online data management unit: used to provide open interface services to meet the data access needs of different user roles.

[0013] Preferably, the data management module adds a data asset health assessment model, which is used to quantitatively assess the operational status of global data assets based on a multi-dimensional indicator system including data quality, service readiness, and security risks.

[0014] Preferably, the special report generation unit uses graph neural network-driven intelligent narrative report generation technology to encode metadata knowledge graphs through GNN models, thereby achieving in-depth analysis and automated natural language description of the complex correlation features between corrosion events, environmental impacts, and facility status.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention efficiently integrates data resources from multiple fields such as meteorology, ecological environment, geology, hydrodynamics, and facility corrosion. It employs a dynamic data cleaning and auditing method based on Flink and Hive for integrated big data streaming and batch processing, constructing a comprehensive and systematic environmental corrosion monitoring data resource pool. This effectively solves the "data silo" phenomenon in existing environmental corrosion monitoring data management systems, achieving data fusion and unification. 2. This invention also provides data resource sharing and services through an online data management unit. Through established data reporting and distribution templates, combined with reporting standards, it generates summary reports for submission and on-demand distribution. 3. This invention integrates databases and operating systems... The invention achieves domestic substitution in core aspects such as system, development framework, and security encryption. Simultaneously, through open-source components, it constructs an independently controllable environmental corrosion data management platform, achieving complete domestic self-reliance and controllability. 4. This invention, through the anomaly detection capability of the cleaning and auditing unit, introduces advanced statistical methods such as the Grubbs test and unsupervised machine learning models, accurately capturing complex nonlinear anomalies and subtle pattern deviations caused by sensor malfunctions, significantly improving the depth of data cleaning and the effective scoring of data. 5. The invention's special report generation unit, through a graph-to-text generation model based on graph neural networks (GNN), achieves a leap from simple numerical filling to intelligent narrative based on knowledge graph reasoning. The report content possesses stronger in-depth insight, logical reasoning, and natural language description capabilities, greatly enhancing the decision-making value of the report. 6. This invention, by introducing a multi-dimensional data asset health assessment model into the data asset management unit, expands the asset assessment system from basic quality to lineage integrity, service timeliness, etc., providing a scientific and quantitative basis for the operation and maintenance of data assets. Attached Figure Description

[0016] Figure 1 This is a block diagram of the overall structure and functional modules of an environmental corrosion data storage and management system according to the present invention; Figure 2 This is a diagram illustrating the overall architecture of an environmental corrosion data storage and management system according to the present invention. Figure 3 This is a technical architecture diagram of the overall structure of an environmental corrosion data storage and management system according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] This invention is based on a B / S architecture and uses the Spring Cloud Alibaba microservice framework and Spring Cloud Gateway service gateway.

[0019] Please see Figure 1-3 This invention provides a technical solution: an environmental corrosion data storage and management system, comprising: Data acquisition and input module: Used for online input and offline filling of multi-source heterogeneous data such as meteorology, ecological environment, geology, hydrodynamics, and facility corrosion, to realize the timed online acquisition of data from various databases and the offline import of multi-source data, and to ensure the accuracy and integrity of the data; The online data acquisition feature automatically retrieves data from external data sources (such as weather stations, environmental monitoring stations, corrosion data centers, etc.) in real time or at set intervals through preset data interfaces or data connections. This ensures the timeliness and accuracy of the data, reduces manual intervention, improves work efficiency, and supports the calling of vector data, raster data, tabular data, image data, video data, and text data files. Offline data entry utilizes data entry templates provided by the system for various topics. To ensure the accuracy and reliability of the offline data, a multi-level review mechanism has been added to the offline data entry unit, including user self-inspection, manual review and marking, and expert review processes. This ensures that manually entered or corrected data has a high degree of reliability before entering the cleaning process. Taking the "Facility Corrosion" topic as an example, the entry template is shown below:

[0020] Manually input or upload data files, and the system can automatically recognize input fields, comments, data information, and create tables and store them in the database. It can handle data that cannot be obtained in real time or requires manual verification, ensuring the accuracy and integrity of the data. It supports input of multiple source data types such as Excel and CSV. The aforementioned online guidance specifically includes: Data Source Management Unit: This unit simultaneously connects to multiple databases from various sources. It enables online data collection from each database through designed data access interfaces, supporting well-known domestic databases such as DM, Kingbase, ShenTong, Highgo, and Youxuan. DM database connection example: Driver class name: dm.jdbc.driver.DmDriver URL format: jdbc:dm: / / <host> : <port> / <database_name> <host>: The IP address or hostname of the DM database server.

[0021] <port>The listening port for the DM database is 5236 by default.

[0022] <database_name> : The name of the database instance to connect to.

[0023] Authentication method: username and password.

[0024] Assume that the DM database is running on 192.168.1.100, port 5236, database name TESTDB, username SYSDBA, and password 123456.

[0025] Configure snippets in the DataX configuration: { "job":{ "setting":{ "speed":{ "byte":1048576 } }, "content":[ { "reader":{ "name":"rdbmsreader", "parameter":{ "username":"SYSDBA", "password":"123456", "connection":[ { "jdbcUrl":[ "jdbc:dm: / / 192.168.1.100:5236 / TESTDB" ], "table":[ "your_table_name ] } ], "splitPk":"id", "where":"1=1" } }, "writer":{ } } ] } }

[0026] The data source management unit adopts a web-based system method that is compatible with multiple databases (including common domestic and foreign databases) based on the DataX component. The web-based system provides a unified user interface and real-time monitoring and management functions. The system based on the DataX component enables dynamic configuration of database connection information and data exchange and conversion between different databases. DataX has a universal data type system that converts data from the source database into its internal type when reading data, and then converts the internal type back into the corresponding type of the target database when writing data to the target database.

[0027] Taking MySQL's datetime function and the DM database as examples: MySQL's datetime type: typically represents date and time, accurate to the second.

[0028] DM Database's date and time types: DM Database's date and time types include DATETIME (date and time, accurate to the second) and TIMESTAMP (time stamp, accurate to the nanosecond), etc.

[0029] When DataX reads datetime data from MySQL, the mysqlreader plugin converts it to DataX's internal DATE type (which can represent both date and time). Then, when the dmwriter plugin writes this data to the DM database, it converts the DataX internal DATE type to the DM database's DATETIME or TIMESTAMP type. Typically, DataX chooses the best matching or most compatible target type for conversion.

[0030] Syntactic differences handling: The SQL syntax does indeed differ between different databases, for example: Pagination syntax: MySQL uses LIMITOFFSET, Oracle uses ROWNUM, and DM database usually also supports LIMITOFFSET or simulates it through ROWNUM.

[0031] Function names: Date functions, string functions, etc., may have different names and usages in different databases.

[0032] DDL syntax: The syntax for DDL statements such as creating tables and modifying tables also differs.

[0033] DataX handles syntax differences in the following ways: Reader plugin: Requires users to provide SQL statements that conform to the source database syntax.

[0034] Writer plugin: Internally generates DML statements that conform to the target database syntax.

[0035] Data type conversion: Internally, it implements a mapping between general types and database-specific types; Task Management Unit: Used to manage projects, collect templates, construct lead tasks, and arrange visual collection tasks, and set lead task parameters (including task name, scheduling cycle, etc.) to achieve scheduled collection of database data.

[0036] Data governance module: To ensure the integrity, standardization, authenticity and reliability of multi-source heterogeneous data storage, it is used to judge data anomalies and missing data, and to identify correct, suspicious, erroneous and missing data, realize data integration and traceability, and support the full life cycle management and quality improvement of data; The data governance module includes: The data cleaning and review unit addresses errors, redundancy, or missing data in monitoring data. It designs data cleaning methods through components, providing data access, cleaning, and transformation, and designs cleaning rules for automated review. A rule base is established, storing various cleaning and review rules. Each rule is associated with one or more "label expressions" or "label conditions." When a new dataset enters the system, it reads its accompanying data labels and then matches these labels against the corresponding rules in the rule base to initially verify the rationality and validity of the collected data. Cleaning rule example: Null value filling: Replaces null values ​​in a specific field with the default value or "unknown".

[0037] Data standardization: Convert "Yes / No" and "Y / N" to "1 / 0".

[0038] Unit conversion: Convert Celsius to Fahrenheit.

[0039] Outlier smoothing: Smoothing outbursts in sensor readings.

[0040] Example of review rules and expressions: Numerical rules: For example, "Field A must be a positive integer".

[0041] Formula rules: For example, if there are three fields A, B, and C, it must satisfy either A>B or A+B=C.

[0042] Formatting rules: such as date format, phone number format, string encoding, etc.

[0043] Example of an automatic formula validation rule: Range check: Checks whether a numeric field is between a predefined minimum and maximum value.

[0044] Incrementality check: Checks whether the value of a sequence field always increases (or decreases). For example: the timestamp of the current record must be greater than the timestamp of the previous record.

[0045] Continuity check: This checks for missing or skipped data in the sequence. For example, the difference between the current temperature and the previous temperature should not exceed a maximum threshold to prevent sudden data changes.

[0046] Extreme value testing: Identifying abnormally high or low extreme values ​​in a dataset that may indicate measurement errors or anomalous events.

[0047] To further improve the depth and accuracy of cleaning, the cleaning review unit introduces multi-model anomaly detection technology: Enhanced statistical testing: For key time-series indicators such as corrosion rate, ambient temperature, and relative humidity, Grubbs test and Kolmonov-Smirnov test have been added to accurately identify outliers in the dataset and verify whether the distribution of the collected data is within the expected statistical range.

[0048] Machine learning anomaly detection: Introducing unsupervised learning algorithms, such as autoencoders or IsolationForest, to learn the underlying patterns of normal time series data and identify complex pattern deviations or subtle sensor malfunctions in the time series.

[0049] The cleaning and auditing unit adopts a dynamic processing method for multi-source heterogeneous data that integrates Flink and Hive for big data stream and batch processing. Data sets are described by data tags, enabling the system to dynamically adapt to different datasets during operation. Appropriate cleaning and auditing rules are loaded according to the dataset tags to complete the data processing. Flink serves as the stream processing engine for real-time data stream processing and cleaning, Hive serves as the batch processing tool for historical data auditing and analysis, and Kafka performs real-time message synchronization. By adopting the concept of stream and batch processing, the system can process both bounded and unbounded stream data simultaneously. Metadata Management Unit: Used for metadata maintenance and processing. After selecting the processing version (metadata / post-approval / model output), data type (internal / external), processing time, status, table name, etc., metadata details can be queried. Lineage and correlation analysis of metadata information can be performed, establishing upstream and downstream relationships, constructing data resource links, and comprehensively displaying the overall data flow. The technical architecture for upstream and downstream relationships is as follows: Metadata collector: Uses DataX to connect to a relational database via a JDBC driver to obtain table structure and view definitions.

[0050] Metadata storage: Using a domestic database, the collected metadata is stored, including information such as tables, columns, views, stored procedures, ETL tasks, reports, and the relationships between them.

[0051] Bloodline analysis engine: Using Apache Atlas, it parses and analyzes data stored in the metadata store, constructs bloodline graphs, and provides a query interface.

[0052] Knowledge Graph Construction: The metadata management unit employs a graph data model, extending metadata management to the construction of a corrosion monitoring knowledge graph. Nodes in the knowledge graph include facilities, sensors, environmental parameters, corrosion events, and quality assessment reports; edges represent physical connections, logical dependencies, corrosion drivers, etc. This graph forms the basis for intelligent report generation in the next section.

[0053] Front-end display: The kinship chart is displayed in a visual way, allowing users to intuitively see the flow and transformation process of the data.

[0054] Data Model: The core of using a graph data model lies in representing the relationship between data assets (tables, columns, files, metrics, etc.) and data operations / transformations (ETL tasks, SQL statements, scripts, etc.).

[0055] Simplified map data model example: Suppose we have an ETL task that reads data from source_table_A and source_table_B, transforms it, and writes it to target_table_C.

[0056] Table: source_table_A (Attributes: name, database, schema, etc.) Table: source_table_B (Attributes: name, database, schema, etc.) ETL_Job:Job_XYZ (Attributes: Name, Responsible Person, Execution Time, etc.) Table: target_table_C (Attributes: name, database, schema, etc.) process (ETL_Job:Job_XYZ)-[READS_FROM]->(Table:source_table_A) (ETL_Job:Job_XYZ)-[READS_FROM]->(Table:source_table_B) (ETL_Job:Job_XYZ)-[WRITES_TO]->(Table:target_table_C) Data quality standard calculation formula: Validity rate = ((number of data records that meet all validity rules) / (total number of data records)) * 100%; Similarly, the formulas for missing detection rate and duplication rate are similar. In a specific embodiment, if the "phone number" field of 50 out of 1000 user records is empty, then the missing detection rate for phone numbers is 5%. In 1000 order records, if 20 records are found to be duplicates based on the order number, then the order duplication rate is 2%.

[0057] Data Standard Unit: Used to define data quality standards such as validity rate, missing rate, and duplication rate, and to assign weights to the data quality standard items, as shown in the table below:

[0058] Data Quality Unit: Used to monitor, display, alert on anomalies, and generate data quality analysis reports according to data quality standards. Data management module: used to add, delete, modify, query and retrieve various monitoring data, and generate data compilation and data asset statistical reports; The data management module includes: Data editing unit: Used to support the editing of time-series data and attribute data, enabling the modification, querying, and maintenance management of monitoring data, analysis results, and attribute data; Data query unit: It supports map query, attribute query, collection time query, analysis report query, data product query, image query and video query. Among them, map query can use point selection, box selection and polygon selection to query the monitoring point name, station number, monitoring data, etc. within the corresponding area. Data Asset Management Unit: Used for multi-dimensional monitoring and statistics of global data assets, including asset scale, asset object assessment, and asset cluster distribution; Data Asset Health Assessment: The data asset management unit adds a multi-dimensional data asset health assessment model, expanding the assessment system to a comprehensive "health" index, including: (1) Basic quality: maintain integrity, validity and uniqueness indicators, and include the proportion of outliers identified through advanced anomaly detection.

[0059] (2) Structure and Relationship: Increase the integrity of bloodline links and the cross-topic relevance index based on knowledge graph (to measure the strength of the synergistic effect of data in corrosion prediction).

[0060] (3) Value and service: Add timeliness indicators (such as real-time stream acquisition delay) and service readiness (such as API call success rate and downstream trend prediction model data input matching degree).

[0061] Thematic Data Statistics Unit: Used to perform data statistics on various thematic data indicators, supporting data statistics by time, space, attributes and other dimensions, and displaying the statistical results in charts; Thematic Report Generation Unit: This unit is used to design analysis report templates based on the reporting standards for each thematic topic, and to statistically analyze the elements of each thematic topic by year, month, and day. The statistical values ​​and results are automatically filled in the corresponding positions of the template. The thematic report generation unit includes an introduction, overview, monitoring work overview, ecological environment status, conclusions and recommendations, and appendices. The monitoring work overview includes monitoring content, evaluation content, monitoring quality control, etc. The ecological environment status includes monitoring results and current status evaluation of each element, monitoring and evaluation methods, etc. Intelligent narrative report generation technology: The thematic report generation unit adopts a graph-to-text (G2T) model based on graph neural networks (GNN) to automate the process from data insights to natural language narrative. (1) Graph encoder: Using a GNN model, such as a graph attention network (GAT) or a graph Transformer architecture, as an encoder, the knowledge graph constructed by the metadata management unit is deeply encoded to extract the complex correlation features between corrosion events, environmental impacts and facility status.

[0062] (2) Sequence decoding: Using a sequence decoder based on Transformer or RNN / LSTM, the graph structure vector encoded by GNN is transformed into coherent and logical natural language text.

[0063] (3) Automatic narrative: The model can automatically generate report narratives with logical reasoning, especially automatically analyze and write conclusions and recommendations, as well as chapters that require comprehensive analysis, such as ecological environment status assessment.

[0064] Online Data Management Unit: This unit provides various interface services for different user roles to facilitate the diverse data access needs of third-party users. The online data management unit includes: Data transmission: Data reporting and distribution templates are developed for each specialty. Data is filled in according to the template content to realize the summary reporting and on-demand distribution of data; Service Management: Through data service management and form service management, external entities can access data and charts via API. Data service management publishes data as corresponding data services and adds data descriptions. By selecting data tags, time ranges, service names, and approving data services, the service interface is exposed to the outside world, enabling data querying and previewing in web page format. Correspondingly, data service queries can view access request information (such as request examples, response examples, access IP and port), and access service management displays all data service access requests. The predicted results (such as predicted future corrosion rates and remaining useful life (RUL)) are incorporated into the data management module as high-value data assets. To efficiently support downstream early warning systems' millisecond-level access to high-frequency time-series data, the API interface provided by the service management unit preferably uses a low-latency time-series database (TSDB) for data storage and retrieval.

[0065] In this invention, service management generates corresponding data tags based on service type and name; it generates visualizations in different chart formats based on service type and theme to display service content and service type for different users; and it provides external data interface service management, including data interface document management and external interface monitoring, displaying information such as request frequency, request time period, and concurrency in a visual manner to ensure the stability and security of external data services.

[0066] By integrating multiple modules such as data acquisition and reception, data governance, and data management, the system can efficiently integrate data resources from various fields, including meteorology, ecology and environment, geology, hydrodynamics, and facility corrosion. It has overcome key technical challenges in the storage, governance, and management of multi-source heterogeneous data, constructing a comprehensive and systematic environmental corrosion monitoring data resource pool. This invention, tailored to the characteristics of domestically developed platforms, effectively solves the problems of "data silos" and resource fragmentation in existing environmental corrosion monitoring data management systems, achieving data fusion and unification, and laying a solid foundation for the efficient utilization of data.

[0067] In terms of data governance, this invention uses methods such as data cleaning, auditing, metadata management, data standardization, and data quality monitoring to identify data anomalies and missing data. It also identifies correct, questionable, erroneous, and missing data, enabling data integration and traceability, effectively improving the integrity, standardization, authenticity, and reliability of the data. This not only supports the entire data lifecycle management but also significantly improves data quality, providing strong support for subsequent data analysis and decision-making.

[0068] This invention provides data resource sharing and services through an online data management unit, offering various interface services for different user roles to meet the diverse data access needs of third-party users. Furthermore, based on established data reporting and distribution templates and filling standards, it can generate summary reports for submission and on-demand distribution. This further promotes data sharing and exchange, improving data utilization efficiency.

[0069] The present invention also includes a data security module for ensuring system security from multiple aspects such as data storage and system operation, including an access security unit, a backup security unit, a log management unit, and a data hardening unit.

[0070] In summary, by integrating multi-source heterogeneous data, establishing data sharing and exchange standards, and strengthening data governance and management functions, this system has significantly improved the management level and utilization efficiency of environmental corrosion monitoring data, providing strong data support and decision support for regional ecological environmental protection and development.

[0071] This invention also provides steps for an environmental corrosion data storage and management system applied to a domestically developed platform. The system adopts a microservice architecture, and all components are selected to ensure system smooth operation and stable deployment. Figure 3 As shown: 1) The application system under the B / S architecture is compatible with three web browsers: Firefox, 360, and Qi An Xin. 2) This invention uses the Tongxin operating system, the MINIO object storage system, the Elasticsearch search server, the DataEase data visualization engine, the Kafka message middleware, and the DolphinScheduler ETL scheduling tool, among other software environments. 3) Databases closely related to data security use domestically produced databases, and sensitive data is encrypted and stored (AES algorithm). The database system has the characteristics of high security, high availability and ease of use, and is a relational database that meets national security standards; 4) The backend uses the Spring Boot development framework, the distributed storage uses HDFS, the frontend uses the Vue development framework, and Nacos provides a reliable distributed coordination and configuration management platform to solve some key problems in distributed systems, such as distributed locks, leader election, configuration management, and distributed synchronization. 5) Because the system needs to access multi-source heterogeneous data, the DataX component, an open-source component from Alibaba, was selected to enable data access from various data sources. 6) Currently, there are many data indicators for each topic. Considering the need for rapid response in subsequent data analysis, we use the domestic open-source BI component DataEase. Through custom data connections and chart selection, statistical analysis can be quickly achieved. 7) At the same time, due to the large amount of data and the variety of data types, the Atlas component was selected to realize data governance and metadata management, which helps to understand the structure, relationship and purpose of data resources, and lays a solid foundation for the establishment of subsequent data standards.

[0072] This invention achieves domestic substitution in core areas such as database, operating system, development framework, and security encryption. At the same time, it builds an independent and controllable environmental corrosion data management platform through open source components, achieving complete domestic independence and control, and effectively reducing dependence on foreign technologies.

[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0074] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.< / port> < / host> < / port> < / host>

Claims

1. An environmental corrosion data storage and management system, characterized in that: include: Data acquisition and reception module, data governance module, and data management module; The data acquisition and reception module is used to integrate multi-source heterogeneous data from meteorological, ecological environment, geological, hydrodynamic and facility corrosion fields to realize online data reception and offline data entry. The data governance module includes a cleaning and auditing unit that uses a dynamic processing method based on Flink and Hive, combined with Grubbs test and machine learning unsupervised model for data cleaning and auditing, and a metadata management unit that constructs a corrosion monitoring knowledge graph, and constructs the corrosion monitoring knowledge graph through metadata management. The data management module generates intelligent thematic reports based on a graph-to-text generation model using graph neural networks, and employs a multi-dimensional data asset health assessment model for data asset management.

2. The environmental corrosion data storage and management system according to claim 1, characterized in that: The online data transfer specifically involves using the DataX component to establish a database connection and automatically retrieve data from multiple heterogeneous data sources in real time or periodically, ensuring the timeliness and accuracy of the data. The online data transfer specifically includes: Data source management unit: It adopts a web-based system based on DataX components that is compatible with multiple database connections. It is used to connect to multiple databases with multiple data sources simultaneously, and realizes online data collection from each database through the design of data call interfaces. Task Management Unit: Used for project management, data collection templates, task creation, visual data collection task orchestration, setting task parameters, and timed data collection from database tables.

3. The environmental corrosion data storage and management system according to claim 2, characterized in that: The data governance module also includes a data quality standard unit and a data quality unit: The data quality standard unit defines scoring rules for completeness, validity, and uniqueness; The data quality unit performs data quality monitoring, anomaly alarms, and quality analysis report generation based on the data quality standard unit.

4. The environmental corrosion data storage and management system according to claim 3, characterized in that: The metadata management unit: DataX is used to collect table structure and view definition metadata, enabling unified storage of metadata. By parsing metadata relationships using Apache Atlas, a data lineage graph can be constructed. A knowledge graph for corrosion monitoring is constructed using a graph data model. The nodes of the knowledge graph include monitoring facilities, environmental parameters, corrosion events, and quality reports, while the edges include physical associations, logical dependencies, and corrosion driving factors.

5. The environmental corrosion data storage and management system according to claim 4, characterized in that: The data governance module's cleaning and auditing unit has a rule base that stores cleaning rules for null value filling, data standardization, unit conversion, and outlier smoothing, as well as auditing rules for numerical rules, formula rules, and format rules. The machine learning unsupervised model is an autoencoder or IsolationForest, used to identify time-series data pattern deviations and sensor malfunctions.

6. The environmental corrosion data storage and management system according to claim 5, characterized in that: The multi-dimensional data asset health assessment model includes: a basic quality dimension including the proportion of outliers; a structural and relational dimension including the integrity of the lineage and the cross-topic relevance index; and a value and service dimension including real-time stream acquisition latency and API call success rate.

7. The environmental corrosion data storage and management system according to claim 6, characterized in that: In the process of generating the intelligent thematic report, the graph neural network model uses GAT or graph Transformer as the encoder to extract knowledge graph association features, and generates conclusions and suggestions with logical reasoning and ecological environment status assessment through Transformer or RNN / LSTM decoder.

8. The environmental corrosion data storage and management system according to claim 7, characterized in that: The data management module includes: Data editing unit: used to support the editing of time-series and attribute data; Data query unit: used to support multiple query methods, including map query; Data Asset Management Unit: Used for monitoring and statistical analysis of global data assets; Special topic data statistics unit: used to support multi-dimensional data statistics; Special Topic Report Generation Unit: Used to generate analysis reports based on the reporting standards for each topic; Online data management unit: used to provide open interface services to meet the data access needs of different user roles.

Citation Information

Patent Citations

  • Data management system for coexistence of multiple complex business lines of Internet

    CN114756563A

  • Ground surface abnormity early warning text generation method based on expression knowledge graph

    CN117708346A

  • Building processing quality safety monitoring system

    CN120931434A

  • Domain large model geological survey report generation method based on knowledge graph

    CN121145871A

  • Computer big data information processing system

    CN121255895A