Multi-source data management system for ecological environment monitoring
Through the multi-layered architecture of the multi-source data management system, multiple problems in data collection, transmission, storage, and analysis in the ecological environment monitoring system have been solved, achieving efficient, secure, and intelligent data management, supporting multi-terminal collaboration and hierarchical control, and improving the efficiency of environmental governance decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA NAT ENVIRONMENTAL MONITORING CENT
- Filing Date
- 2026-02-04
- Publication Date
- 2026-04-21
AI Technical Summary
Existing ecological and environmental monitoring systems lack sufficient data collection dimensions and flexibility, are deficient in transmission reliability and security, have low storage and retrieval efficiency, weak data processing and intelligent analysis capabilities, and imperfect management and auditing mechanisms, making it difficult to meet diverse needs.
A multi-source data management system is adopted, including a multi-source data acquisition layer, a data encryption transmission layer, a data hierarchical storage layer, and a data intelligent analysis layer. Through standardized interfaces, data and control information can be exchanged, enabling accurate acquisition, secure transmission, efficient storage, and intelligent analysis of multi-source data.
It has improved the efficiency and security of ecological and environmental monitoring data collection, enhanced the efficiency of data storage and retrieval, strengthened the intelligent analysis capabilities of data processing, realized multi-terminal collaborative management and hierarchical control, and supported scientific decision-making and precise supervision.
Smart Images

Figure CN121907889A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ecological environment monitoring, and in particular to a multi-source data management system for ecological environment monitoring. Background Technology
[0002] In the field of ecological and environmental monitoring, existing data management systems generally adopt a traditional architecture of "distributed collection + centralized storage." The corresponding data management method is as follows: the data collection stage uses a single type of ground sensor to collect time-series monitoring data at fixed intervals; then, the time-series monitoring data is transmitted to a central server via a 4G network or wired network; finally, a relational database is used to store the time-series monitoring data. Functionally, these systems only provide basic data query and simple chart display services; system operation and maintenance management adopts a single administrator model, lacking a multi-role collaboration mechanism.
[0003] However, existing data management systems are ill-suited to the diverse needs of complex ecological environment monitoring in practical applications, and suffer from numerous prominent problems: (1) Insufficient data collection dimensions and flexibility: The data collection equipment is limited to a single type or a few types of ground sensors, which cannot achieve the collaborative collection of multi-dimensional data such as air, water quality, soil, meteorology and remote sensing, and the data coverage is limited; moreover, the sampling frequency is fixed and cannot be dynamically adjusted according to changes in environmental parameters or sudden environmental events, which can easily lead to the omission of key data collection.
[0004] (2) Lack of transmission reliability and security: Over-reliance on traditional mobile communication or wired networks can easily lead to transmission interruptions and data loss in remote and weak signal areas. At the same time, the data transmission process often uses weak encryption methods, and some scenarios even lack end-to-end integrity verification mechanisms, posing a security risk of data tampering and leakage.
[0005] (3) Low storage and retrieval efficiency: Relational databases have poor adaptability to the storage of large-scale time-series monitoring data and are difficult to support the efficient management of unstructured data such as remote sensing images and videos; the backup strategy is mostly periodic full backup, which not only occupies a lot of storage resources, but also causes the data recovery time to be too long, affecting emergency use.
[0006] (4) Weak data processing and intelligent analysis capabilities: It can only achieve statistical summary and simple graphical display of data, lacks effective functions for abnormal data identification, multi-source data fusion and environmental trend prediction, cannot deeply explore the value of data, and is difficult to support scientific decision-making and precise supervision.
[0007] (5) Inadequate management and auditing mechanisms: lack of multi-terminal adaptation capabilities, unable to meet the operational needs of different scenarios; at the same time, lack of fine-grained permission control, and no comprehensive operation log traceability system has been established, making it difficult to achieve hierarchical management and accountability.
[0008] Therefore, there is an urgent need to develop a new multi-source data management system for ecological and environmental monitoring to solve one or more of the above problems. Summary of the Invention
[0009] The purpose of this application is to provide a multi-source data management system for ecological and environmental monitoring, which can realize accurate collection, secure transmission, efficient storage, intelligent analysis and hierarchical management of multi-source data, and comprehensively improve the value of monitoring data and the efficiency of environmental governance decision-making.
[0010] To achieve the above objectives, this application provides a multi-source data management system for ecological environment monitoring, comprising: a multi-source data acquisition layer, a data encryption transmission layer, a data hierarchical storage layer, a data intelligent analysis layer, and a multi-terminal control layer. Each layer interacts with control information through standardized interfaces. Specifically, the multi-source data acquisition layer acquires multi-source data, dynamically adjusts and standardizes the sampling frequency of the multi-source data to obtain standardized multi-source data. The multi-source data includes at least: air monitoring data, water quality monitoring data, soil monitoring data, meteorological monitoring data, satellite remote sensing data, and UAV aerial photography data. The data encryption transmission layer receives the standardized multi-source data, performs effective data filtering and redundancy removal on the standardized multi-source data to obtain filtered data, and encrypts the filtered data to obtain... Encrypted data is securely pushed to a hierarchical data storage layer via transmission networks adapted to different scenarios. The hierarchical data storage layer receives and decrypts the encrypted data, categorizes and stores it according to the data type of the decrypted multi-source data, and implements hierarchical management, multi-layer backup, and index optimization for the stored multi-source data, achieving efficient storage and rapid retrieval. The intelligent data analysis layer calls upon the stored multi-source data to sequentially complete data preprocessing and fusion, anomaly detection and multi-level early warning, and short / medium-term trend prediction, generating intuitive multi-source data analysis results. The multi-terminal management layer connects to the intuitive multi-source data analysis results, supports multi-terminal collaborative access and real-time early warning push, and achieves hierarchical management and operation traceability of the entire multi-source data process through fine-grained access control and tamper-proof audit logs.
[0011] As described above, the multi-source data acquisition layer includes: sensor arrays and edge computing nodes deployed in each monitoring area, remote sensing access units, and data format units; wherein, the sensor arrays are used to collect multi-source monitoring data, which includes at least: air monitoring data, water quality monitoring data, soil monitoring data, and meteorological monitoring data; the remote sensing access units are used to receive satellite remote sensing data and UAV aerial photography data; the data format units adopt a unified time-series data model to standardize the multi-source data and obtain standardized multi-source data.
[0012] As mentioned above, the multi-source data acquisition layer also includes a dynamic acquisition control module, which is integrated into the edge computing node and is used to regulate the adaptive acquisition of multi-source data.
[0013] As mentioned above, the multi-source data acquisition layer also includes a sensor health prediction module. The sensor health prediction module is deployed in parallel with the dynamic acquisition control module to proactively predict sensor failures / attenuation before the sensor array performs multi-source monitoring data acquisition.
[0014] As described above, the sub-steps by which the sensor health prediction module proactively predicts sensor failure / attenuation before the sensor array performs multi-source monitoring data acquisition are as follows: E1: Obtain historical data from each sensor in the sensor array from the multi-source data acquisition layer. This historical data includes at least: calibration records within a preset first time period, real-time data within a preset second time period, and concurrent data from a reference sensor at the same location; E2: Calculate three-dimensional index scores based on the historical data. These scores include: calibration effectiveness score, data stability score, and error rate score; E3: Calculate sensor health based on the three-dimensional index scores and analyze the sensor health using a preset health threshold. If the sensor health is greater than or equal to the health threshold, the sensor array performs multi-source monitoring data acquisition; if the sensor health is less than the health threshold, a backup sensor is automatically switched, and a maintenance work order is triggered. The expression for sensor health is: ; in, For sensor health; These are the weighting coefficients. ; The calibration validity score; Score the data stability. The score is the error rate.
[0015] As described above, the data encryption transmission layer includes: a transmission network module, an edge preprocessing module, a security encryption module, and a transmission assurance module. The transmission network module has a transmission network strategy: in cellular network coverage areas, edge computing nodes prioritize 5G or 4G networks for low-latency data transmission; in scenarios without cellular coverage or low-power monitoring, edge computing nodes automatically switch to BeiDou short message, LoRa, or satellite short message as a fallback transmission solution; and encrypted data is reported to the cloud platform through a segmented retransmission strategy and the transmission network strategy. The edge preprocessing module is deployed within the edge computing node, equipped with a lightweight containerized service. After receiving standardized multi-source data, it performs preprocessing operations on the standardized multi-source data locally on the edge computing node, obtains filtered data, and sends it to the security encryption module. The module includes a fully encrypted module and a secure encryption module. The secure transmission mechanism utilizes a public key infrastructure (PKI) to perform bidirectional authentication between the sensor array and edge computing nodes in the multi-source data acquisition layer. If the sensor array is valid, authentication is successful, and the sensor array sends standardized multi-source data to the edge preprocessing module. If the sensor array is invalid, authentication fails, and the data transmission link is blocked. A hybrid encryption scheme is used to perform hybrid encryption on the filtered data to obtain initial encrypted data. This hybrid encryption scheme involves negotiating and generating a symmetric key during the session establishment phase using an asymmetric encryption algorithm. Subsequent filtered data is encrypted using this symmetric key to obtain initial encrypted data. A message authentication code is also appended to the initial encrypted data to obtain the final encrypted data. The secure encryption module's transmission link synchronously supports TLS. 1.2 / 1.3 Protocol and Encrypted Data Segmentation and Retransmission Strategy; The encrypted data and its segmentation and retransmission strategy are sent to the transmission network module; Transmission Guarantee Module: Encrypted data is transmitted to the hierarchical data storage layer through multiple collaborative mechanisms. Specifically, a message queue with acknowledgment mechanism or a preset transmission protocol is used to obtain real-time feedback on the cloud platform's reception status of the encrypted data segments; For encrypted data segments that have completed encryption and are awaiting transmission, persistent caching is performed in the local storage medium of the edge computing node; A breakpoint resumption mechanism is configured so that when network connectivity is restored, the segmented data of the encrypted data that was not successfully transmitted is automatically identified and filtered out, and only the targeted retransmission of this segmented data is triggered, without repeating the segmented data that has been successfully delivered to the cloud platform.
[0016] As described above, the data hierarchical storage layer includes: a storage media adaptation module, a data hierarchical management and control module, a backup and recovery management module, and an index retrieval optimization module. The storage media adaptation module includes a hybrid storage architecture consisting of a time-series database, distributed object storage, and a relational database. The time-series database stores high-frequency time-series data from multiple sources. The distributed object storage stores unstructured data from multiple sources, including at least satellite remote sensing data and UAV aerial photography data. The relational database stores metadata and access control data. The data hierarchical management and control module divides the multi-source data into three levels: basic monitoring data, key monitoring data, and anomaly monitoring data, based on parameter type, trigger threshold, and access frequency. For key monitoring data and anomaly monitoring data... The management and control methods for routine monitoring data are as follows: configure a multi-replica high-availability storage mechanism and deploy an off-site disaster recovery solution to synchronize data to off-site backup storage nodes in real time; the management and control methods for basic monitoring data are as follows: adopt a single-replica or dual-replica low-redundancy storage strategy and do not deploy off-site disaster recovery; the backup and recovery management module: build a multi-layer backup system by combining incremental backup, snapshot technology and segmented recovery technology to ensure the recoverability of multi-source data; the index retrieval optimization module: establish time partition indexes and tag indexes for the time series database to improve the query efficiency of time series data; establish metadata indexes for unstructured data in distributed object storage, and the index dimensions of the metadata indexes should at least include geographical coordinates, collection time and spectral bands, supporting fast retrieval based on spatial range and time range.
[0017] As described above, the data intelligence analysis layer includes: a preprocessing and fusion unit, an anomaly detection and early warning unit, a trend prediction unit, a model training and online update unit, and a visualization and reporting unit. Specifically, the preprocessing and fusion unit: calls upon multi-source data stored in the hierarchical data storage layer, performs advanced cleaning operations on the called multi-source data in the cloud, and obtains cleaned data; the advanced cleaning operations include at least inter-sensor calibration and sensor drift compensation; and uses a machine learning regression / fusion model to fuse the cleaned data to obtain fused data. The anomaly detection and early warning unit: constructs a deep learning-based time-series anomaly detection model, combines it with a rule engine to execute multi-level early warnings on the fused data, and obtains anomaly warning information; When an alert is triggered, alert evidence is recorded simultaneously. The alert evidence includes at least the relevant data segment, triggering rule, and confidence level. Trend prediction unit: uses time series model and deep learning model to predict short-term / medium-term trends based on the fused data, generating a prediction report including confidence intervals. Model training and online update unit: based on the fused data and anomaly alert information, provides two training modes for machine learning regression / fusion models and deep learning-based time series anomaly detection models: offline batch training or online incremental learning, to obtain optimized models. Saves model version information and metadata of the training dataset. Visualization and reporting unit: transforms anomaly alert information and prediction reports into intuitive multi-source data analysis results.
[0018] As described above, the multi-terminal control layer includes: a multi-terminal adaptation unit, a permission management unit, and an audit and logging unit. The multi-terminal adaptation unit supports access from multiple terminals, including PC management terminals, mobile apps, and WeChat mini-programs, allowing users to access intuitive multi-source data analysis results. The permission management unit uses the RBAC permission model, supporting fine-grained access control policies based on resources, actions, and time. Administrators can configure the mapping relationship between roles and permissions, and the system integrates multi-factor authentication to ensure access security. The audit and logging unit records the entire user operation log, including at least login, data access, configuration changes, and model retraining. The user operation log uses structured storage and is tamper-proof. It supports log retention and audit evidence export according to preset policies for compliance verification and operation tracing.
[0019] As shown above, the standardized interface is either a RESTful API or an asynchronous message queue. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0021] Figure 1 This is a schematic diagram of one embodiment of a multi-source data management system for ecological environment monitoring. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] like Figure 1 As shown, this application provides a multi-source data management system for ecological environment monitoring, including: a multi-source data acquisition layer, a data encryption transmission layer, a data hierarchical storage layer, a data intelligent analysis layer, and a multi-terminal control layer. Each layer realizes data and control information interaction through standardized interfaces.
[0024] The multi-source data acquisition layer collects multi-source data, dynamically adjusts and standardizes the sampling frequency of the multi-source data to obtain standardized multi-source data. The multi-source data includes at least: air monitoring data, water quality monitoring data, soil monitoring data, meteorological monitoring data, satellite remote sensing data, and UAV aerial photography data.
[0025] Data encryption transmission layer: Receives standardized multi-source data, performs effective data filtering and redundancy removal on the standardized multi-source data to obtain filtered data; encrypts the filtered data to obtain encrypted data; and securely pushes the encrypted data to the data hierarchical storage layer through a transmission network adapted to different scenarios.
[0026] Data hierarchical storage layer: Receives encrypted data and decrypts it, classifies and stores the decrypted multi-source data according to its data type, implements hierarchical management, multi-layer backup and index optimization for the stored multi-source data, and achieves efficient storage and fast retrieval of multi-source data.
[0027] Data intelligence analysis layer: It calls upon stored multi-source data to sequentially complete data preprocessing and fusion, anomaly detection and multi-level early warning, and short / medium-term trend prediction, generating intuitive multi-source data analysis results.
[0028] Multi-terminal control layer: It connects to intuitive multi-source data analysis results, supports multi-terminal collaborative access and real-time early warning push, and realizes hierarchical management and operation traceability of multi-source data throughout the entire process through fine-grained access control and tamper-proof audit logs.
[0029] Furthermore, the standardized interface is a predefined communication protocol interface used to realize data transmission and control command interaction between various layers (i.e., multi-source data acquisition layer, data encryption transmission layer, data hierarchical storage layer, data intelligent analysis layer, and multi-terminal management and control layer). The interface type is determined according to the data interaction scenario and transmission requirements between the corresponding layers. In this application, it is preferred to be a RESTful API (representational state transition style application programming interface) or an asynchronous message queue (such as MQTT (message queue telemetry transmission)) to adapt to the low-latency transmission of massive time-series monitoring data, but it is not limited to RESTful API or asynchronous message queue.
[0030] Furthermore, the multi-source data acquisition layer includes: sensor arrays and edge computing nodes deployed in each monitoring area, remote sensing access units, and data format units.
[0031] The sensor array is used to collect multi-source monitoring data, which includes at least: air monitoring data, water quality monitoring data, soil monitoring data and meteorological monitoring data; the sensor types in the sensor array include at least: air monitoring terminal, water quality monitoring terminal, soil monitoring terminal and meteorological monitoring terminal.
[0032] The remote sensing access unit is used to receive satellite remote sensing data and drone aerial photography data.
[0033] The data format unit adopts a unified time-series data model to standardize multi-source data, obtain standardized multi-source data, and realize standardized storage and cross-level interaction of data from different sources and types. It records the data source ID, geographic coordinates, sampling time, sampling frequency, sensor status, and calibration information of monitoring data collected by various sensors. Among them, the identifier dimension of the unified time-series data model corresponds to the data source ID and geographic coordinates, the time dimension of the unified time-series data model corresponds to the sampling time and sampling frequency, and the data dimension of the unified time-series data model corresponds to the sensor status and calibration information.
[0034] Specifically, the calibration parameters, sampling accuracy, and upper limit of error for each type of sensor are set according to the national / industry standards for the corresponding monitoring indicators, the actual monitoring scenario requirements, and the purpose of data application.
[0035] Air monitoring terminals are used to collect air monitoring data, which includes at least: PM2.5 monitoring values, PM10 monitoring values, and so on. monitoring values Monitoring values and The monitored values.
[0036] The water quality monitoring terminal is used to collect water quality monitoring data, which includes at least the following: pH value, dissolved oxygen value, COD value, and ammonia nitrogen value.
[0037] Soil monitoring terminals are used to collect soil monitoring data, which includes at least the following: soil moisture monitoring values, soil temperature monitoring values, and soil heavy metal ion concentration monitoring values.
[0038] The meteorological monitoring terminal is used to collect meteorological monitoring data, which includes at least the following: wind speed monitoring value, wind direction monitoring value, precipitation monitoring value for a preset time period, air temperature monitoring value, and air humidity monitoring value.
[0039] The remote sensing access unit supports the acquisition and coordinate registration of multiple data formats, including at least GeoTIFF and image tiles.
[0040] A unified time-series data model can be flexibly selected based on the actual monitoring data scale, system deployment cost, and functional requirements. For example, in the scenario of massive IoT time-series data, open-source general models can be selected, including the InfluxDB Line Protocol data model (adapted for efficient writing and querying) and the Prometheus time-series data model (supporting multi-dimensional label classification analysis). If it is necessary to adapt to the lightweight deployment requirements of this system, a custom lightweight time-series data model can be selected, which has a core structure of three layers: time dimension (collection timestamp, sampling period), identification dimension (monitoring point ID, indicator type), and data dimension (collected value, accuracy, status identifier).
[0041] Furthermore, the multi-source data acquisition layer also includes a dynamic acquisition control module, which is integrated into the edge computing node and is used to regulate the adaptive acquisition of multi-source data.
[0042] Specifically, the dynamic acquisition and control module is a configurable rule engine based on thresholds and rates of change.
[0043] Furthermore, the configurable rule engine based on thresholds and rates of change allows users to flexibly configure judgment rules according to actual ecological environment monitoring needs. Preferably, the rule configuration items of the configurable rule engine based on thresholds and rates of change include, but are not limited to, monitoring indicator thresholds, rate of change per unit time thresholds, sampling frequency multiples after triggering, a list of associated monitoring indicators, and the duration of high-frequency sampling.
[0044] Specifically, an exemplary configuration rule for the configurable rule engine based on thresholds and rates of change is as follows: If the sliding window rate of change of the xth monitoring data is greater than X% (X% can be flexibly configured according to the monitoring scenario, such as 10% for soil moisture monitoring and 15% for air pollutant monitoring), or if the real-time value of the monitoring data is within a preset threshold trigger range (for example, the preset threshold trigger range is 80%~95% of the upper limit of the threshold), then the sampling interval of the monitoring data is adjusted from the default period T1 to a short sampling period T2 (T1 and T2 can be flexibly configured according to the monitoring indicator type; in normal scenarios, T1 is instantiated as 1 hour and T2 as 5 minutes; in high-precision scientific research monitoring scenarios, T1 can be configured as 30 minutes and T2 as 1 minute). At the same time, the triggering reason and corresponding timestamp of this sampling interval adjustment are recorded synchronously to achieve full data traceability. Through the configurable rule engine based on thresholds and rates of change, the technical pain point of the traditional fixed-frequency sampling mode being unable to respond to sudden environmental changes in a timely manner is effectively solved. While ensuring the complete collection of key monitoring data, unnecessary redundant data accumulation is avoided, significantly improving data collection efficiency and data validity.
[0045] Furthermore, the multi-source data acquisition layer also includes a sensor health prediction module. The sensor health prediction module is deployed in parallel with the dynamic acquisition control module. It is used to proactively predict sensor failures / attenuation before the sensor array performs multi-source monitoring data acquisition, replacing the passive processing mode of abnormal data and ensuring the continuity of multi-source data.
[0046] Furthermore, the sub-steps of proactively predicting sensor failure / attenuation before the sensor array performs multi-source monitoring data acquisition via the sensor health prediction module are as follows: E1: Obtain historical data from each sensor in the sensor array from the multi-source data acquisition layer. The historical data includes at least: calibration records within a preset first time period, real-time data within a preset second time period, and synchronous data from a reference sensor at the same location.
[0047] Specifically, the range of the preset first and second time periods can be flexibly set according to the actual scenario. For example, the calibration records of the previous 30 days starting from the current collection time, the real-time data of the previous 24 hours starting from the current collection time, and the data of the reference sensor at the same location during the same period (reusing existing data).
[0048] E2: Calculates three-dimensional index scores based on historical data. The three-dimensional index scores include: calibration effectiveness score, data stability score, and error rate score.
[0049] Furthermore, the expression for calculating the three-dimensional indicators based on historical data is as follows: ; in, The calibration validity score; This is the normalization coefficient for the average verification deviation; This represents the average calibration deviation of the calibration records within the preset first time period.
[0050] The specific value can be flexibly set according to actual needs. This application prefers the following: .
[0051] ; in, Score the data stability. The normalization coefficient for the ratio of standard deviations; The standard deviation of the real-time data within the preset second time period; The standard deviation of the real-time data within the preset third time period is defined as the second time period, which is part of the third time period. This represents the ratio of standard deviations.
[0052] Specifically, the scope of the third time period is set according to the actual scenario. In this application, it is preferred to be the first three months starting from the time of this data collection. The specific value can be flexibly set according to actual needs. This application prefers the following: .
[0053] ; in, Score the error rate; This is the normalization coefficient for the average error rate; The average error rate between real-time data and the reference sensor during the preset second time period.
[0054] Specifically, The specific value can be flexibly set according to actual needs. This application prefers the following: .
[0055] E3: Calculate sensor health based on three-dimensional index scores, and analyze sensor health using preset health thresholds. If sensor health is greater than or equal to the health threshold, the sensor array performs multi-source monitoring data acquisition; if sensor health is less than the health threshold, the backup sensor is automatically switched and a maintenance work order is triggered.
[0056] Furthermore, the expression for sensor health is: ; in, For sensor health; These are the weighting coefficients. .
[0057] Specifically, The specific value is set according to the actual scenario. Preferably, in this application: .
[0058] Furthermore, the data encryption transmission layer includes: a transmission network module, an edge preprocessing module, a security encryption module, and a transmission guarantee module.
[0059] Among them, the transmission network module: is provided with a transmission network policy, and the transmission network policy is: in the area covered by the cellular network, the edge computing node preferentially selects the 5G or 4G network to complete low-latency data transmission; in the scenario of no cellular coverage or low-power monitoring, the edge computing node automatically switches to Beidou short message, LoRa (Long Range Radio) or satellite short message as the fallback transmission solution; the encrypted data is reported to the cloud platform through the segmented retransmission policy of the encrypted data and the transmission network policy.
[0060] The edge preprocessing module: is deployed in the edge computing node, equipped with a lightweight containerized service. After receiving the standardized multi-source data, it completes the preprocessing operation of the standardized multi-source data locally in the edge computing node, obtains the screened data and sends it to the security encryption module; among them, the preprocessing operation at least includes data cleaning, redundant data deduplication, format standardization and unification, multi-source data time series alignment, and preliminary anomaly detection; the screened data includes: effective monitoring data and anomaly warning data.
[0061] The security encryption module: is provided with a secure transmission mechanism, and the secure transmission mechanism is: based on the Public Key Infrastructure (PKI), it completes the two-way authentication of the identities of the sensor array in the multi-source data acquisition layer and the edge computing node. If the sensor array (i.e., the access device) is legal, it means the authentication is successful, and the sensor array sends the standardized multi-source data to the edge preprocessing module. If the sensor array (i.e., the access device) is illegal, it means the authentication fails, and the data transmission link is blocked; a hybrid encryption scheme is used to perform hybrid encryption on the effective monitoring data and the anomaly warning data to obtain the initial encrypted data; among them, the hybrid encryption scheme is: through an asymmetric encryption algorithm (such as RSA or ECC), a symmetric key (such as AES-256) is negotiated and generated in the session establishment stage, and subsequent screened data is encrypted using this symmetric key to obtain the initial encrypted data; at the same time, a Message Authentication Code (HMAC-SHA256) is attached to the initial encrypted data to obtain the encrypted data; the transmission link synchronization supporting the TLS 1.2 / 1.3 protocol and the segmented retransmission policy of the encrypted data is provided for the security encryption module; the encrypted data and the segmented retransmission policy of the encrypted data are sent to the transmission network module.
[0062] Transport Assurance Module: Transmits encrypted data to the data hierarchical storage layer through multiple collaborative mechanisms to ensure the reliable transmission of encrypted data and resume interrupted transmissions. The specific implementation method is as follows: Adopt a message queue with a message confirmation mechanism or a preset transmission protocol (such as the MQTT QoS protocol, the AMQP protocol) to obtain the feedback on the reception status of the segmented data of the encrypted data from the cloud platform in real time; For the segmented data of the encrypted data that has been encrypted and is waiting to be transmitted, perform a persistent caching operation in the local storage medium of the edge computing node to avoid the risk of data loss caused by sudden network interruptions, abnormal downtime of edge computing node devices, etc.; Configure a mechanism for resuming interrupted transmissions. When the network resumes connectivity, automatically identify and filter out the segmented data of the encrypted data that has not been successfully transmitted, and only trigger the targeted retransmission of this part of the segmented data, without repeating the transmission of the segmented data that has already been successfully delivered to the cloud platform.
[0063] Specifically, reporting the effective monitoring data and abnormal alarm data to the cloud platform through the transmission network module can significantly reduce the transmission ratio of invalid data and save the bandwidth resources of the cloud platform. The specific type of lightweight containerized service is flexibly selected according to the actual monitoring scenario. For example, adopt Docker containers or container orchestration technology. Adopt a hybrid encryption scheme to further improve the security and reliability of data transmission. Attach a message authentication code (HMAC-SHA256) to the encrypted initial data to verify the integrity of the encrypted data through the message authentication code.
[0064] In the transmission link assurance mechanism supporting the security encryption module of this application, the segmented retransmission strategy for encrypted data refers to: For the effective monitoring data and abnormal alarm data blocks after hybrid encryption and message authentication code verification, divide them into several data segments with unique serial number identifiers according to a preset byte size threshold; Only record the transmission status of each segment during the transmission process. If the cloud platform feedbacks that a segment with a certain serial number is lost or the verification fails, only retransmit the abnormal segment, without retransmitting the complete data block; After all segments are successfully received, the cloud platform reassembles them into a complete encrypted data block according to the serial number, and then decrypts and performs subsequent processing.
[0065] Furthermore, each edge computing node can real-time identify the network reachability status. Each edge computing node autonomously selects an appropriate transmission channel according to the network reachability status according to the transmission network strategy, and synchronously records the key transmission information. Among them, the network reachability status at least includes: the connectivity status of the cellular network (5G or 4G), the availability status of Beidou short messages, the availability status of LoRa, and the availability status of satellite short messages; The key transmission information at least includes: the type of transmission channel, the switching time, and the transmission status.
[0066] Furthermore, the data hierarchical storage layer includes: a storage medium adaptation module, a data hierarchical control module, a backup and recovery management module, and an index retrieval optimization module.
[0067] The storage media adaptation module includes a hybrid storage architecture consisting of a time-series database, a distributed object storage system, and a relational database. The time-series database (e.g., InfluxDB, TimescaleDB) is used to store high-frequency time-series data from multiple sources. The distributed object storage system (e.g., HDFS, Ceph, S3 compatible) is used to store unstructured data from multiple sources, including at least satellite remote sensing data and UAV aerial photography data. The relational database is used to store metadata and access control data.
[0068] The data hierarchical management module classifies multi-source data into three levels: basic monitoring data, critical monitoring data, and abnormal monitoring data, based on parameter type, trigger threshold, and access frequency. The management approach for critical and abnormal monitoring data involves configuring a multi-replica high-availability storage mechanism and deploying an off-site disaster recovery solution to synchronize data to an off-site backup storage node in real time. The management approach for basic monitoring data involves adopting a low-redundancy storage strategy with single or dual replicas, without deploying off-site disaster recovery, to balance storage costs and data availability.
[0069] Backup and recovery management module: Combining incremental backup, snapshot technology and segmented recovery technology to build a multi-layer backup system to ensure the recoverability of multi-source data.
[0070] Index retrieval optimization module: Creates time-partitioned indexes and tag indexes for time-series databases to improve the query efficiency of time-series data; creates metadata indexes for unstructured data in distributed object storage. The index dimensions of the metadata indexes include at least geographic coordinates, collection time and spectral bands, supporting fast retrieval based on spatial and temporal ranges.
[0071] Specifically, high-frequency time-series data refers to continuously monitored parameter data collected at fixed time intervals such as seconds or minutes, with each data point having a timestamp. Examples include real-time temperature and humidity, water pH, and atmospheric particulate matter concentration data in environmental monitoring. Multi-replica high-availability storage mechanisms refer to storing multiple identical copies of the same data on multiple independent physical nodes in a distributed storage cluster. When any storage node fails and becomes inaccessible, the system can automatically switch to a copy on another working storage node to provide data service, thus ensuring 24 / 7 uninterrupted data availability. Off-site disaster recovery solutions refer to deploying backup storage systems in different regions (such as different cities or provinces) physically isolated from the local data center. Critical data is replicated to this backup system through real-time synchronization or scheduled backups. When the local storage system is paralyzed due to extreme circumstances such as natural disasters or equipment failures, data service can be quickly restored through the off-site backup system.
[0072] Furthermore, based on parameter type, trigger threshold, and access frequency, multi-source data is divided into three levels: basic monitoring data, key monitoring data, and abnormal monitoring data. The three-level classification rule is as follows: Level 1 (basic monitoring data): monitoring data with parameters of routine monitoring items, no explicit threshold trigger requirements, and a weekly access frequency of less than 10 times; Level 2 (key monitoring data): monitoring data with parameters of core monitoring items (such as drinking water source water quality indicators), trigger thresholds related to public safety standards, and a weekly access frequency of more than 50 times; Level 3 (abnormal monitoring data): all monitoring data with parameter values exceeding the preset safety threshold (such as water quality heavy metal content exceeding the standard).
[0073] Incremental backup, snapshot, and segmented recovery technologies are all mature existing technologies in the data backup field. Multi-layered backup systems can be built based on existing backup software or native database functions. A multi-layered backup system ensures the recoverability of monitoring data, preventing permanent data loss due to storage node failures, accidental data deletion, natural disasters, etc., while reducing the cost of storage resources and network bandwidth during the backup process. In practice, a tiered backup cycle of daily incremental backups, weekly differential backups, and monthly full snapshots is adopted. For critical and abnormal monitoring data, additional data is synchronized to local backup nodes and off-site cold backup systems. Segmented recovery technology divides the backup data into independent data segments according to time range or data type, achieving rapid and accurate recovery of the target data.
[0074] Furthermore, the data intelligence analysis layer includes: a preprocessing and fusion unit, an anomaly detection and early warning unit, a trend prediction unit, a model training and online update unit, and a visualization and reporting unit.
[0075] The preprocessing and fusion unit calls upon multi-source data stored in the hierarchical data storage layer, performs advanced cleaning operations on the called multi-source data in the cloud, and obtains cleaned data. The advanced cleaning operations include at least inter-sensor calibration and sensor drift compensation. The cleaned data is fused using machine learning regression / fusion models (e.g., random forest, XGBoost) to obtain fused data, thereby improving data integrity and consistency.
[0076] Anomaly Detection and Early Warning Unit: Constructs a time-series anomaly detection model based on deep learning, and performs multi-level early warnings on the fused data in conjunction with a rule engine to obtain anomaly early warning information; when an early warning is triggered, early warning evidence is recorded simultaneously. The early warning evidence includes at least the relevant data segment, the triggering rule, and the confidence level, which is used to support the tracing of the early warning results.
[0077] Trend prediction unit: Using time series models (e.g., ARIMA, Prophet) and deep learning models (e.g., LSTM, Transformer time series variants) based on fused data, it performs short-term / medium-term trend predictions and generates prediction reports containing confidence intervals to assist in decision-making.
[0078] Model Training and Online Update Unit: Based on the fused data and anomaly warning information, it provides two training modes for machine learning regression / fusion models and deep learning-based time-series anomaly detection models: offline batch training or online incremental learning, to obtain optimized models; it saves model version information and metadata of the training dataset for model backtracking and auditing.
[0079] Visualization and Reporting Unit: Transforms anomaly warning information and forecast reports into intuitive multi-source data analysis results.
[0080] Specifically, the visualization and reporting unit provides various visualization formats, including at least line charts, heatmaps, spatiotemporal overlay satellite imagery, and thematic maps; it supports custom dashboard configurations and timed / event-driven report generation (output format: PDF / HTML) for intuitive display and distribution of data results.
[0081] Specifically, time-series anomaly detection models based on deep learning can be built using existing technologies, such as LSTM-AE and Seq2Seq prediction residual threshold detection.
[0082] The rule engine can be implemented using existing technology without additional development. This application requires the customization of multi-level early warning rule sets for monitoring scenarios. For example: Level 1 early warning: A single monitoring indicator exceeds the safety threshold by 5% for 5 consecutive minutes, and the model detection confidence is ≥80%; Level 2 early warning: ≥2 monitoring indicators exceed the safety threshold at the same time, and the model detection confidence is ≥90%; Level 3 early warning: A key monitoring indicator (such as the heavy metal content in drinking water sources) instantly exceeds the safety threshold by 2 times, with a confidence of ≥95%.
[0083] Short-term forecasts refer to the changing trends of monitoring indicators within hours to 7 days, while medium-term forecasts refer to the changing trends of monitoring indicators within 7 days to 3 months. The specific duration can be flexibly configured according to the actual monitoring scenario requirements. The two models can be used as follows: the time series model and the deep learning model can run independently to output prediction results, or they can be weighted and fused together to improve accuracy.
[0084] Furthermore, the data intelligence analysis layer also includes a cross-regional data collaborative analysis module, which is used to realize cross-regional collaborative anomaly detection and linkage trend prediction based on multi-regional and multi-source data stored in the data hierarchical storage layer, and support governance decisions in scenarios such as cross-regional pollution diffusion and watershed water quality linkage.
[0085] Furthermore, the cross-regional data collaborative analysis module performs the following steps: T1: Retrieves multi-source storage data from multiple regions from the data hierarchical storage layer. The multi-source storage data includes at least: geographical coordinates of each region, multi-source data of each region, and data collection timestamps.
[0086] Specifically, the multi-source data for each region includes: air monitoring data, water quality monitoring data, soil monitoring data, meteorological monitoring data, satellite remote sensing data, and drone aerial photography data.
[0087] T2: Based on preset core monitoring indicators, core indicators are selected from multi-source stored data, and feature vectors for each region are constructed based on the core indicators.
[0088] Specifically, the details of selecting core monitoring indicators can be flexibly set according to the actual scenario. For example, selecting core monitoring indicators includes: mandatory indicators, optional indicators, and exclusion rules; among which, mandatory indicators are selected from multi-source stored data according to the core monitoring items specified in national / industry standards (such as PM2.5 for air quality, etc.). For water quality, pH and dissolved oxygen are mandatory; for soil, heavy metal concentration is mandatory. Optional indicators: core indicators can be dynamically added according to the monitoring scenario (e.g., "basin flow" for cross-basin scenarios, "industrial emission point data" for urban agglomeration scenarios). Removal rules: indicators with a data missing rate >30% in a preset period (e.g., the last 30 days) will be removed to avoid interference from low-quality data.
[0089] Among them, the Feature vectors of the region The expression is: , The number of core indicators.
[0090] Existing technologies can be used to construct feature vectors for each region based on core indicators, thus unifying data dimensions. For example: data preprocessing tools include Python's Pandas library (which filters core indicator columns using DataFrames and concatenates them into feature vectors) and NumPy library (which converts the filtered core indicators into array-based feature vectors); basic machine learning framework functionalities include Scikit-learn's FeatureUnion class and TensorFlow / PyTorch's tensor construction capabilities (which encapsulate multi-dimensional core indicators into model-processable feature vectors); and time-series data processing tools include InfluxDB and TimescaleDB's query and filtering functions (which first filter out core indicators based on the core monitoring indicators, then export the core indicators as feature vectors). As an example, the feature vector for region A in a cross-basin scenario is: .
[0091] T3: Calculate the regional collaborative weight of the target region in each region based on the feature vector of each region, and take the non-target region with the largest inter-regional collaborative weight as the affected region.
[0092] Furthermore, the expression for calculating the regional collaborative weights of the target regions in each region based on the feature vectors of each region is as follows: ; in, For the first Region and the Regional collaboration weights between regions ; For the first Region and the Geographical distance between regions; for and The absolute value of the Pearson correlation coefficient, ; For the first The feature vector of the region, For the first The feature vector of the region, , The total number of regions; This is the time decay coefficient; For the first The data acquisition timestamps corresponding to the feature vectors of the region; For the first The data acquisition timestamps corresponding to the feature vectors of the region; It is a natural constant.
[0093] Specifically, The larger the value, the more significant the th... Region (affected area) on the first The stronger the influence of the region (target region), the more effective the extraction of metadata. The unit is km. The closer the value is to 1, the better the number of digits. Region and the The more consistent the trends in regional indicators; The specific value can be flexibly set according to the actual scenario, and the preferred value in this application is 0.01.
[0094] T4: Extract single-interval anomalies in the target area and the single-interval anomalies in the affected area from the anomaly warning information obtained from the anomaly detection and warning unit. Correct the single-interval anomalies in the target area based on the regional collaborative weight to obtain collaborative anomalies, and synchronize the collaborative anomalies to the anomaly detection and warning unit.
[0095] Furthermore, based on regional collaborative weighting, the single-interval outliers in the target region are corrected, and the expression for the collaborative outliers is obtained as follows: ; in, For the first Cooperative outliers in the region, the first The region is the target region; For the first Single-interval outliers in the region; For the first Single-interval outliers in the region, the first The area is the region of influence; For the first Region and the Regional collaboration weights between regions; Indicates in Take the maximum value between the two.
[0096] Specifically, , , and The closer it is to 1, the more severe the abnormality.
[0097] Furthermore, when If the value exceeds the preset cross-regional warning threshold, a cross-regional warning is triggered. The specific value of the cross-regional warning threshold can be flexibly set according to the actual scenario; in this application, it is preferably 0.8.
[0098] T5: Based on the trend prediction unit's prediction report and regional collaborative weights, the prediction report for the target area is revised to obtain the revised prediction report.
[0099] Furthermore, the sub-steps for revising the forecast report of the target area based on the forecast report of the trend forecast unit and the regional collaborative weights to obtain the revised forecast report are as follows:
[0100] T51: Extract the predicted change rate of the affected area and the initial predicted value of the target area from the prediction report of the trend prediction unit, and correct the initial predicted value of the single area of the target area based on the regional collaborative weight to obtain the linked predicted value.
[0101] Furthermore, the expression for the linked predicted value is: ; in, For the first Regional linkage prediction values; For the first Initial predicted values for the region; For the first Region and the Regional collaboration weights between regions; For the first Predicted rate of change for the region.
[0102] T52: Extract the confidence interval of the target region from the prediction report of the trend prediction unit, and correct the confidence interval of the target region based on the regional collaborative weight to obtain the corrected confidence interval.
[0103] Furthermore, the revised expression for the confidence interval is: ; in, The confidence interval for the target region; For the first Region and the Regional collaboration weights between regions.
[0104] T53: After correcting the forecast report based on the linked forecast value and the corrected confidence interval, the revised forecast report is obtained.
[0105] Furthermore, the multi-terminal control layer includes: a multi-terminal adaptation unit, a permission management unit, and an audit and logging unit.
[0106] The multi-terminal adaptation unit supports access from multiple terminals, including PC management (Web), mobile APP, and WeChat mini-program, and allows access to intuitive multi-source data analysis results. All terminals interact with the backend system through API (Application Programming Interface) and are configured with a real-time push mechanism (e.g., WebSocket / Push) for timely delivery of alerts and notifications.
[0107] Access Control Unit: Adopts RBAC (Role-Based Access Control) permission model, supports fine-grained access control policies based on resources, actions, and time; administrators can configure the mapping relationship between roles and permissions, and the system integrates multi-factor authentication (MFA) function to ensure access security.
[0108] Audit and Log Unit: Records user operation logs throughout the entire process, including at least login, data access, configuration changes, and model retraining; the user operation logs are stored in a structured manner and are tamper-proof; supports log retention and audit evidence export according to preset policies for compliance verification and operation tracing.
[0109] Specifically, the backend system refers to the core server-side component of the multi-source data management system for ecological environment monitoring used in this application, including all functional modules and hardware resources of the data hierarchical storage layer and the data intelligent analysis layer, which is the core architecture supporting the operation of the entire system.
[0110] Furthermore, the user's full-process operation log adopts a structured storage method, and achieves the function of immutability by writing WORM (Write Once Read Many) or on-chain hash value, but it is not limited to WORM or on-chain hash value.
[0111] The beneficial effects achieved by this application are as follows: (1) The multi-source data management system for ecological environment monitoring in this application breaks through the limitations of traditional single ground sensor acquisition, integrates multi-source data such as air, water quality, soil, meteorology, satellite remote sensing, and UAV aerial photography, and actively predicts faults / attenuation through the sensor health prediction module and the dynamic acquisition control module adaptively adjusts the sampling frequency, which not only expands the data coverage dimension, but also ensures the continuity and accuracy of monitoring data.
[0112] (2) The multi-source data management system for ecological environment monitoring in this application adopts hybrid encryption (i.e., asymmetric negotiation key, symmetric encrypted data and message authentication code), multi-network adaptation (5G / 4G priority, Beidou short message / LoRa backup) and breakpoint resume mechanism, which not only prevents data tampering and leakage risks, but also solves the problem of transmission interruption caused by weak signal in remote areas, so as to achieve secure and reliable data transmission.
[0113] (3) The multi-source data management system for ecological environment monitoring in this application has an innovative hybrid storage architecture, combined with hierarchical management and multi-layer backup, which not only improves data retrieval efficiency but also shortens data recovery time, balancing storage cost and availability.
[0114] (4) The multi-source data management system for ecological and environmental monitoring in this application integrates multi-source data fusion, deep learning anomaly detection, short / medium-term trend prediction and cross-regional collaborative analysis functions to deeply mine data value; at the same time, it supports access from multiple terminals such as PC, mobile APP and WeChat mini program, and combined with RBAC fine-grained permission control and tamper-proof audit logs, it realizes multi-role collaborative management and operation traceability, and adapts to the needs of scientific decision-making and precise supervision.
[0115] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the scope of protection of this application is intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application. Obviously, those skilled in the art can make various alterations and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of protection of this application and its equivalents, this application also intends to include these modifications and variations.
Claims
1. A multi-source data management system for ecological environment monitoring, characterized in that, include: The system consists of a multi-source data acquisition layer, a data encryption transmission layer, a data hierarchical storage layer, a data intelligent analysis layer, and a multi-terminal control layer. Each layer interacts with data and control information through standardized interfaces. The multi-source data acquisition layer collects multi-source data, dynamically adjusts and standardizes the sampling frequency of the multi-source data to obtain standardized multi-source data. The multi-source data includes at least: air monitoring data, water quality monitoring data, soil monitoring data, meteorological monitoring data, satellite remote sensing data, and UAV aerial photography data. Data encryption transmission layer: Receives standardized multi-source data, performs effective data filtering and redundancy removal on the standardized multi-source data to obtain filtered data; encrypts the filtered data to obtain encrypted data; and securely pushes the encrypted data to the data hierarchical storage layer through a transmission network adapted to different scenarios. Data hierarchical storage layer: Receives encrypted data and decrypts it, classifies and stores the decrypted multi-source data according to its data type, implements hierarchical management, multi-layer backup and index optimization for the stored multi-source data, and achieves efficient storage and fast retrieval of multi-source data; Data intelligent analysis layer: It calls upon stored multi-source data to sequentially complete data preprocessing and fusion, anomaly detection and multi-level early warning, and short / medium-term trend prediction, generating intuitive multi-source data analysis results; Multi-terminal control layer: It connects to intuitive multi-source data analysis results, supports multi-terminal collaborative access and real-time early warning push, and realizes hierarchical management and operation traceability of multi-source data throughout the entire process through fine-grained access control and tamper-proof audit logs.
2. The multi-source data management system for ecological environment monitoring according to claim 1, characterized in that, The multi-source data acquisition layer includes: sensor arrays and edge computing nodes deployed in each monitoring area, remote sensing access units, and data format units; The sensor array is used to collect multi-source monitoring data, which includes at least: air monitoring data, water quality monitoring data, soil monitoring data, and meteorological monitoring data. The remote sensing access unit is used to receive satellite remote sensing data and drone aerial photography data; The data format unit adopts a unified time-series data model to standardize multi-source data and obtain standardized multi-source data.
3. The multi-source data management system for ecological environment monitoring according to claim 2, characterized in that, The multi-source data acquisition layer also includes a dynamic acquisition control module, which is integrated into the edge computing node and is used to regulate the adaptive acquisition of multi-source data.
4. The multi-source data management system for ecological environment monitoring according to claim 3, characterized in that, The multi-source data acquisition layer also includes a sensor health prediction module, which is deployed in parallel with the dynamic acquisition control module to proactively predict sensor failures / attenuation before the sensor array performs multi-source monitoring data acquisition.
5. The multi-source data management system for ecological environment monitoring according to claim 4, characterized in that, The sub-steps of proactively predicting sensor failure / attenuation before the sensor array performs multi-source monitoring data acquisition using the sensor health prediction module are as follows: E1: Obtain historical data from each sensor in the sensor array from the multi-source data acquisition layer. The historical data includes at least: calibration records within a preset first time period, real-time data within a preset second time period, and concurrent data from a reference sensor at the same location. E2: Calculates three-dimensional index scores based on historical data. The three-dimensional index scores include: calibration effectiveness score, data stability score, and error rate score. E3: Calculate the sensor health based on the three-dimensional index score, and analyze the sensor health using a preset health threshold. If the sensor health is greater than or equal to the health threshold, the sensor array performs multi-source monitoring data acquisition; if the sensor health is less than the health threshold, automatically switch to the standby sensor and trigger an operation and maintenance work order. Among them, the expression of the sensor health is: ; in, For sensor health; These are the weighting coefficients. ; The calibration validity score; Score the data stability. The score is the error rate.
6. The multi-source data management system for ecological environment monitoring according to claim 5, characterized in that, The data encryption transmission layer includes: a transmission network module, an edge preprocessing module, a security encryption module, and a transmission guarantee module. Among them, the transmission network module: is set with a transmission network policy, and the transmission network policy is: in the area covered by the cellular network, the edge computing node preferentially selects the 5G or 4G network to complete low-latency data transmission; in the scenario of no cellular coverage or low-power monitoring, the edge computing node automatically switches to Beidou short message, LoRa or satellite short message as a fallback transmission solution; report the encrypted data to the cloud platform through the segmented retransmission policy of the encrypted data and the transmission network policy. The edge preprocessing module: is deployed inside the edge computing node, equipped with a lightweight containerized service. After receiving the standardized multi-source data, it completes the preprocessing operation of the standardized multi-source data locally in the edge computing node, obtains the filtered data and sends it to the security encryption module. The security encryption module: is set with a secure transmission mechanism, and the secure transmission mechanism is: based on the public key infrastructure, complete the two-way authentication of the identities of the sensor array and the edge computing node in the multi-source data acquisition layer. If the sensor array is legal, indicating that the authentication is successful, the sensor array sends the standardized multi-source data to the edge preprocessing module. If the sensor array is illegal, indicating that the authentication fails, block the data transmission link; perform hybrid encryption on the filtered data using a hybrid encryption scheme to obtain the initial encrypted data. Among them, the hybrid encryption scheme is: negotiate and generate a symmetric key through an asymmetric encryption algorithm in the session establishment stage, and subsequent filtered data are encrypted using this symmetric key to obtain the initial encrypted data; at the same time, attach a message authentication code to the initial encrypted data to obtain the encrypted data; the transmission link synchronization supporting the security encryption module is compatible with the TLS 1.2 / 1.3 protocol and the segmented retransmission policy of the encrypted data; send the encrypted data and the segmented retransmission policy of the encrypted data to the transmission network module. The transmission guarantee module: transmits the encrypted data to the data hierarchical storage layer through a multiple coordination mechanism. The specific implementation method is: use a message queue with a message confirmation mechanism or a preset transmission protocol to obtain the feedback of the receiving status of the segmented data of the encrypted data from the cloud platform in real time; perform a persistent caching operation on the segmented data of the encrypted data that has been encrypted and waiting to be transmitted in the local storage medium of the edge computing node; configure a breakpoint resumption mechanism. When the network resumes connectivity, automatically identify and filter out the segmented data of the encrypted data that has not been successfully transmitted, and only trigger the targeted retransmission of this part of the segmented data, without repeating the transmission of the segmented data that has been successfully delivered to the cloud platform.
7. The multi-source data management system for ecological environment monitoring according to claim 6, characterized in that, The data hierarchical storage layer includes: a storage medium adaptation module, a data hierarchical control module, a backup and recovery management module, and an index retrieval optimization module. The storage media adaptation module includes a hybrid storage architecture consisting of a time-series database, a distributed object storage system, and a relational database. The time-series database is used to store high-frequency time-series data from multiple sources. The distributed object storage system is used to store unstructured data from multiple sources, including at least satellite remote sensing data and UAV aerial photography data. The relational database is used to store metadata and access control data. The data hierarchical management module classifies multi-source data into three levels: basic monitoring data, critical monitoring data, and abnormal monitoring data, based on parameter type, trigger threshold, and access frequency. The management approach for critical and abnormal monitoring data involves configuring a multi-replica high-availability storage mechanism and deploying an off-site disaster recovery solution to synchronize data to an off-site backup storage node in real time. The management approach for basic monitoring data involves using a single-replica or dual-replica low-redundancy storage strategy without deploying off-site disaster recovery. Backup and recovery management module: Combines incremental backup, snapshot technology and segmented recovery technology to build a multi-layer backup system to ensure the recoverability of multi-source data; Index retrieval optimization module: Creates time-partitioned indexes and tag indexes for time-series databases to improve the query efficiency of time-series data; creates metadata indexes for unstructured data in distributed object storage. The index dimensions of the metadata indexes include at least geographic coordinates, collection time and spectral bands, supporting fast retrieval based on spatial and temporal ranges.
8. The multi-source data management system for ecological environment monitoring according to claim 7, characterized in that, The data intelligence analysis layer includes: a preprocessing and fusion unit, an anomaly detection and early warning unit, a trend prediction unit, a model training and online update unit, and a visualization and reporting unit; The preprocessing and fusion unit calls upon multi-source data stored in the hierarchical data storage layer, performs advanced cleaning operations on the called multi-source data in the cloud, and obtains cleaned data; the advanced cleaning operations include at least inter-sensor calibration and sensor drift compensation; and uses a machine learning regression / fusion model to fuse the cleaned data to obtain fused data. Anomaly Detection and Early Warning Unit: Constructs a time-series anomaly detection model based on deep learning, and performs multi-level early warnings on the fused data in conjunction with a rule engine to obtain anomaly early warning information; when an early warning is triggered, early warning evidence is recorded simultaneously, and the early warning evidence includes at least the relevant data segment, the triggering rule, and the confidence level; Trend Prediction Unit: Uses time series models and deep learning models to predict short-term / medium-term trends based on fused data, and generates a prediction report including confidence intervals; Model Training and Online Update Unit: Based on the fused data and anomaly warning information, it provides two training modes for machine learning regression / fusion models and deep learning-based time-series anomaly detection models: offline batch training or online incremental learning, to obtain optimized models; and saves model version information and metadata of the training dataset. Visualization and Reporting Unit: Transforms anomaly warning information and forecast reports into intuitive multi-source data analysis results.
9. The multi-source data management system for ecological environment monitoring according to claim 8, characterized in that, The multi-terminal control layer includes: a multi-terminal adaptation unit, a permission management unit, and an audit and logging unit; Among them, the multi-terminal adaptation unit is used to support access from multiple terminals such as PC management terminal, mobile APP and WeChat mini program, and to access intuitive multi-source data analysis results; Access Control Unit: Adopts the RBAC access control model, supports fine-grained access control policies based on resources, actions, and time; administrators can configure the mapping relationship between roles and permissions, and the system integrates multi-factor authentication to ensure access security; Audit and Log Unit: Records user operation logs throughout the entire process, including at least login, data access, configuration changes, and model retraining; the user operation logs are stored in a structured manner and are tamper-proof; supports log retention and audit evidence export according to preset policies for compliance verification and operation tracing.
10. The multi-source data management system for ecological environment monitoring according to claim 1, characterized in that, The standardized interface is a RESTful API or an asynchronous message queue.
Citation Information
Patent Citations
Intelligent environment monitoring and early warning system based on multi-sensor fusion
CN120778173A
Information acquisition and transmission system of track control equipment
CN120825504A
Environmental protection management system based on big data service integrated management
CN120911760A
Fixed pollution source automatic monitoring method and system and digital intelligent construction method thereof
CN121303575A
Sensor precision management device and management method using it
KR102743796B1