Ecological environment multi-source big data fusion system oriented to theme and scene application

By constructing a multi-source big data fusion system for the ecological environment, the problems of data silos, poor access compatibility, and lagging updates have been solved. It has achieved efficient unified access and deep integration of multi-source heterogeneous data, improved the application efficiency of data, and supported intelligent ecological environment supervision and emergency decision-making.

CN121615086APending Publication Date: 2026-03-06重庆市生态环境大数据应用中心 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511847652.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing ecological and environmental data systems suffer from data silos, poor data access compatibility, low governance efficiency, insufficient integration depth, weak spatiotemporal analysis capabilities, and outdated update mechanisms, making it difficult for data to support comprehensive analysis and decision-making applications across domains and scenarios.

Method used

Construct an ecological environment multi-source big data fusion system oriented towards themes and scenarios, including a data aggregation layer, a data fusion layer, a data management layer, and a service publishing layer. Employ technologies such as multi-protocol adaptation, edge preprocessing, entity recognition and spatial correlation analysis, distributed indexing, and automatic incremental updates to achieve unified access, intelligent governance, deep fusion, and efficient spatiotemporal indexing of multi-source heterogeneous data.

Benefits of technology

It enables efficient and unified access, intelligent governance and deep integration of multi-source heterogeneous data, provides millisecond-level data response and more than 95% data association accuracy, supports intelligent ecological and environmental supervision and emergency decision-making, and improves the efficiency of data application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121615086A_ABST
    Figure CN121615086A_ABST
Patent Text Reader

Abstract

The invention discloses an ecological environment multi-source big data fusion system oriented to theme and scene application, and relates to the technical field of ecological environment big data processing. Comprising a data collection layer, a data fusion layer, a data management layer and a service publishing layer which interact through a distributed bus, the data collection layer collects, cleans and standardizes ecological environment data; the data fusion layer performs association fusion on the ecological environment theme data by adopting entity recognition and spatial association analysis to generate ecological environment theme fusion data; the data management layer constructs two-dimensional and three-dimensional mixed spatio-temporal indexes for spatio-temporal characteristics in the ecological environment data, and the service publishing layer is used for publishing a scenarized application data set; the system solves the problems of islanding, weak association and slow updating of traditional ecological environment data, can provide millisecond-level data response, and provides high-value data support for ecological environment intelligent supervision and emergency decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological environment monitoring technology, specifically to an ecological environment multi-source big data fusion system oriented towards thematic and scenario applications. Background Technology

[0002] As ecological and environmental protection work advances towards precision and intelligence, the sources of ecological and environmental data have experienced explosive growth, with collection equipment covering various types, resulting in massive amounts of multi-source heterogeneous data. These data differ significantly in format, structure, and spatiotemporal scale, and the construction standards and data specifications of each collection system vary, leading to a serious "information silo" phenomenon. The data lacks effective correlation, making it difficult to support cross-domain and cross-scenario comprehensive analysis and decision-making applications.

[0003] Existing ecological and environmental data processing systems suffer from numerous technical bottlenecks: insufficient data access capabilities and poor compatibility with special data formats; low data governance efficiency, rigid governance rules, and difficulty in adapting to dynamic business changes; insufficient data fusion depth, mostly involving simple field matching, failing to achieve deep fusion at the semantic and business logic levels; weak spatiotemporal analysis capabilities, making it difficult to meet the needs of complex three-dimensional spatial analysis scenarios; and outdated data update mechanisms, often relying on scheduled full synchronization, resulting in delayed data updates and wasted resources. These problems severely restrict the application effectiveness of data in multiple scenarios, necessitating the development of a unified multi-source data access, intelligent governance, thematic fusion, efficient spatiotemporal indexing, and automatic incremental update fusion system. Summary of the Invention

[0004] The purpose of this invention is to provide an ecological environment multi-source big data fusion system oriented towards thematic and scenario applications, which realizes unified collection of multi-source heterogeneous data, intelligent and efficient governance, thematic deep fusion, accurate two-dimensional and three-dimensional spatiotemporal indexing, and automatic incremental updates.

[0005] The basic solution provided by this invention is: an ecological environment multi-source big data fusion system oriented towards thematic and scenario applications, including a data aggregation layer, a data fusion layer, a data management layer, and a service publishing layer that interact through a distributed bus; The data aggregation layer includes a multi-protocol adaptation unit and an edge preprocessing unit. The multi-protocol adaptation unit collects all types of ecological and environmental data through databases, networks, files, and interfaces. The edge preprocessing unit cleans, deduplicates, handles outliers, and standardizes the ecological and environmental data to output high-quality standardized data. The data fusion layer includes a data cache pool and the data governance engine module. The data cache pool and the data governance engine module are communicatively connected. The data governance engine module uses entity recognition and spatial correlation analysis to perform correlation and fusion of ecological and environmental theme data. The generated ecological and environmental theme fusion data is temporarily stored in the data cache pool. The data management layer includes a spatiotemporal index service module and an intelligent update module. The spatiotemporal index service module uses a distributed index architecture to connect to the data cache pool and constructs a two-dimensional and three-dimensional hybrid spatiotemporal index for the spatiotemporal characteristics of ecological and environmental data. The intelligent update module is used for automatic incremental updates and historical version tracking of ecological and environmental theme data. The service publishing layer includes application service interfaces, which are connected to the spatiotemporal index service module and the intelligent update module through a load balancer. The application service interfaces support scenario-based applications and third-party system access through API interfaces, SDK development packages and visualization component libraries.

[0006] The workflow of this system is as follows: 1. Data Acquisition and Preprocessing: The multi-protocol adaptation unit collects all types of ecological and environmental data through various data transmission protocols. The edge preprocessing unit performs localized preliminary processing on the collected data, completing cleaning, deduplication, outlier handling, standardization conversion, and lossless compression of large files. 2. Data association and fusion: The data governance engine module uses an improved BERT-LSTM hybrid model for entity recognition and combines it with spatial association analysis to perform deep association and fusion of ecological and environmental theme data. The fused data is temporarily stored in the data cache pool. 3. Spatiotemporal Index Construction and Data Update: The spatiotemporal index service module constructs a two-dimensional and three-dimensional hybrid spatiotemporal index based on the spatiotemporal characteristics of the data, while the intelligent update module realizes automatic incremental updates and historical version tracking of the data; 4. Service Release and Application Support: The service release layer provides support for scenario-based applications and third-party system access through various forms of interfaces and component libraries, enabling efficient data application.

[0007] The advantages of this solution are as follows: This invention constructs a complete multi-source big data fusion processing architecture, realizing unified access, intelligent governance, deep fusion, efficient indexing, and automatic updates of multi-source heterogeneous data; through the collaborative work of various functional units, it solves the pain points of traditional ecological and environmental data being isolated, weakly correlated, and slow to update; the system can provide millisecond-level data response, over 95% data correlation accuracy, and full-process data quality control, providing high-value data support for intelligent ecological and environmental supervision and emergency decision-making, and has broad application prospects and practical value.

[0008] Furthermore, the multi-protocol adaptation unit includes data transmission protocols that match HTTP / HTTPS, MQTT, Kafka, FTP / SFTP, and JDBC / ODBC, and reserves a custom protocol extension interface.

[0009] The compatibility of data with different transmission methods solves the problems of single data access and poor compatibility in traditional systems; the reserved custom protocol extension interface enables the system to flexibly adapt to newly added special data transmission protocols in the future, enhancing the system's scalability and life cycle, and meeting the needs of the continuous development of ecological and environmental data acquisition technology.

[0010] Furthermore, the edge preprocessing unit includes a lightweight ETL tool embedded in the embedded computing chip to perform localized preliminary processing of ecological environment data, including key field integrity verification, data validity verification, and lossless compression processing of large video and image files.

[0011] By embedding lightweight ETL tools into embedded computing chips, local preliminary data processing is achieved, eliminating the need to transmit all raw data to a backend server for further processing. This reduces bandwidth consumption and latency during data transmission, improving data processing efficiency. Key field integrity checks and data validity verification can filter out some invalid data in advance, reducing the backend processing pressure. Lossless compression processing of large video and image files reduces data storage space while ensuring data integrity, and simultaneously improves the speed of data transmission and subsequent processing.

[0012] Furthermore, the data aggregation layer also includes a dynamic rule engine unit, which adopts a visual drag-and-drop rule configuration interface to customize input data cleaning rules, including numerical range rules, logical verification rules, and correlation verification rules.

[0013] The visual drag-and-drop rule configuration interface lowers the technical threshold for rule configuration, allowing users to customize data cleaning rules without the need for technical professionals to write complex code.

[0014] Furthermore, the data aggregation layer also includes a data quality assessment unit, which evaluates the completeness, accuracy, consistency and timeliness of ecological and environmental data through a data quality assessment model. When the assessment score is lower than the threshold, the ecological and environmental data to be assessed is sent to the manual review process.

[0015] The data quality assessment model comprehensively evaluates the completeness, accuracy, consistency, and timeliness of the data, ensuring that the data entering the subsequent process meets the quality standards and providing a reliable data foundation for subsequent data integration and application. The assessment threshold and manual review process are set to perform secondary verification and correction on unqualified data, avoiding the impact of low-quality data on the system analysis results, and further improving the reliability and usability of the data.

[0016] Furthermore, the comprehensive score calculation of the data quality assessment model satisfies the following formula: ; Where S is the overall data quality score. Scores are awarded for completeness, accuracy, consistency, and timeliness. For the weights corresponding to the basic dimensions, ; For the scores of N custom dimensions, For the weights corresponding to the custom dimensions, Quantitative assessment methods make data quality evaluation more objective and accurate, avoiding the subjectivity and ambiguity of traditional qualitative evaluation, and facilitating the system to quantitatively manage and control data quality.

[0017] Furthermore, the entity recognition in the data governance engine module is implemented using an improved BERT-LSTM hybrid model, satisfying the following logic:

[0018] Where y is the predicted label sequence, and A is the transition probability matrix. Indicates from the label Move to label The probability, Z(x) represents the probability of the label y_i corresponding to the i-th word in the model output, and Z(x) is the normalization factor.

[0019] The improved BERT-LSTM hybrid model combines the powerful semantic understanding capabilities of the BERT model with the advantages of the LSTM model in processing sequence data, enabling more accurate identification of entity information in ecological and environmental subject data and improving the accuracy of entity recognition.

[0020] Furthermore, the spatiotemporal index service module adopts a composite index structure of R-tree, quadtree, and time axis. The three-dimensional spatial data of elevation information uses R-tree index, and the three-dimensional spatial data is divided into layers by minimum boundary rectangle. The two-dimensional planar data uses quadtree index, and the planar space is recursively divided into four quadrants according to longitude and latitude. The time dimension is associated with the spatial index through the time axis index. During the index construction process, a spatiotemporal layering strategy and a hot data caching mechanism are used to achieve dynamic optimization, and the response time of spatiotemporal combination query is controlled within 500 milliseconds.

[0021] A composite index structure of R-tree, quadtree, and time axis is adopted, and appropriate indexing methods are used for three-dimensional spatial data (elevation information) and two-dimensional planar data respectively. At the same time, the time dimension is associated with the spatial index, so as to achieve comprehensive coverage and efficient indexing of the spatiotemporal characteristics of ecological and environmental data.

[0022] Furthermore, in the construction of the R-tree index for the three-dimensional spatial data, the three-dimensional range of the minimum bounding rectangle (MBR) satisfies the following relationship:

[0023] in, , These represent the minimum and maximum coordinate values ​​of the three-dimensional spatial data along the x-axis, respectively. , These are the minimum and maximum coordinate values ​​along the y-axis, respectively. , These are the minimum and maximum coordinate values ​​in the z-axis direction, respectively.

[0024] By precisely defining the boundaries of 3D spatial data using the maximum and minimum coordinate values ​​in the x, y, and z axes, a clear and standardized calculation basis is provided for the hierarchical division and efficient indexing of 3D spatial data. This ensures the accuracy and rationality of the 3D spatial data index construction and further enhances the system's ability to process and query 3D spatial data. Attached Figure Description

[0025] Figure 1 This is a diagram illustrating the overall architecture of the multi-source big data fusion system for the ecological environment. Figure 2 This is a detailed functional module structure diagram of the data aggregation layer. Detailed Implementation

[0026] The following detailed description illustrates the specific implementation method: like Figure 1 As shown in the figure, this embodiment provides an ecological environment multi-source big data fusion system oriented towards theme and scenario applications, including a data aggregation layer, a data fusion layer, a data management layer and a service publishing layer that interact through a distributed bus.

[0027] like Figure 2 As shown, the data aggregation layer includes a multi-protocol adaptation unit, an edge preprocessing unit, a dynamic rule engine unit, and a data quality assessment unit.

[0028] The multi-protocol adaptation unit includes data transmission protocols that match HTTP / HTTPS, MQTT, Kafka, FTP / SFTP, and JDBC / ODBC, and reserves interfaces for custom protocol extensions; it collects all types of ecological environment data through databases, networks, files, and interfaces.

[0029] Developed using Java, the multi-protocol adaptation unit adopts the Spring Cloud framework and integrates client components for protocols such as HTTP / HTTPS, MQTT, Kafka, FTP / SFTP, and JDBC / ODBC. It allows users to specify the source, protocol type, and connection parameters of the data to be collected through configuration files, enabling unified collection of data from different protocols. It also reserves a custom protocol extension interface, supporting the addition of new protocols through a plug-in approach. Users can develop custom protocol plugins and integrate them into the system.

[0030] The edge preprocessing unit includes a lightweight ETL tool embedded in the embedded computing chip to perform localized preliminary processing of ecological and environmental data, including key field integrity verification, data validity verification, and lossless compression processing of large video and image files. It also cleans, deduplicates, handles outliers, and standardizes ecological and environmental data to output high-quality standardized data.

[0031] Deploy lightweight ETL tools in the Linux operating system of the embedded industrial computer, and write data processing scripts to perform integrity checks on key fields; check whether core fields in the data are missing, such as monitoring time, monitoring location, monitoring index values, and verify data validity, judging the validity of the data based on the reasonable range of monitoring indicators, such as PM2.5 concentration cannot be negative; remove duplicate data based on the unique data identifier field; and convert data of different formats to a unified data format of the system, such as unifying the date format to "YYYY-MM-DD HH:MM:SS". Compress video files using H.265 encoding and image files using JPEG2000 encoding.

[0032] The dynamic rule engine unit uses a visual drag-and-drop rule configuration interface to customize input data cleaning rules, including numerical range rules, logical verification rules, and correlation verification rules.

[0033] A visual drag-and-drop rule configuration interface is developed using Vue.js, allowing users to drag and drop rule components such as numerical range components, logical judgment components, and association verification components. Data cleaning rules are built, and after configuration, corresponding rule scripts are automatically generated and stored in the rule database. During the data cleaning process, the system reads the rule scripts and performs targeted data cleaning. For example, users can configure a rule for PM2.5 concentration ranges of 0-1000 μg / m³, and the system will automatically filter out PM2.5 concentration data outside this range.

[0034] The data quality assessment unit evaluates the completeness, accuracy, consistency, and timeliness of ecological and environmental data using a data quality assessment model. When the assessment score is below the threshold, the ecological and environmental data being assessed is sent to the manual review process.

[0035] A data quality assessment model was developed based on Python. The comprehensive score of the data quality assessment model is calculated according to the following formula: ; Where S is the overall data quality score. Scores are awarded for completeness, accuracy, consistency, and timeliness. For the weights corresponding to the basic dimensions, ; For the scores of N custom dimensions, For the weights corresponding to the custom dimensions, .

[0036] The scores for the four basic dimensions—completeness, accuracy, consistency, and timeliness—are calculated as follows: Completeness Score S1: Completeness Score = (Number of core fields actually included / Total number of core fields that should be included) × 100; Accuracy Score S2: Accuracy Score = (Number of valid data points / Total number of data points) × 100; Consistency Score S3: Consistency Score = (Number of data entries with consistent format / Total number of data entries) × 100; Timeliness Score S4: If the difference between the data collection time and the data arrival time in the system is ≤1 hour, the timeliness score is 100; if 1 hour < difference ≤2 hours, the timeliness score is 80; if the difference >2 hours, the timeliness score is 60.

[0037] The weights w1, w2, w3, and w4 of the basic dimensions can be adjusted according to business needs. The default settings are w1=0.25, w2=0.3, w3=0.15, and w4=0.1, with a total of 0.8.

[0038] Custom dimensions can be added by users based on specific business scenarios. For example, for the "data acquisition device trustworthiness" dimension, users can set trustworthiness weights for different devices and calculate the score S for this dimension based on device trustworthiness. j Custom dimension weights w j The total score is 0.2. When the overall data quality score S < 80, the system automatically sends the batch of data to the manual review platform for manual review and correction by staff.

[0039] The data fusion layer includes a data cache pool and the data governance engine module. The data cache pool and the data governance engine module are communicatively connected. The data governance engine module uses entity recognition and spatial correlation analysis to correlate and fuse ecological and environmental data under the same theme, and the generated ecological and environmental fusion data is temporarily stored in the data cache pool.

[0040] A distributed data cache pool is built using Redis to temporarily store pre-processed and cleaned raw data as well as merged thematic data, with a data expiration policy set. For example, merged data is cached for 72 hours, and raw data is cached for 24 hours to avoid excessive storage space being occupied by cached data.

[0041] Entity recognition is performed based on an improved BERT-LSTM hybrid model. The model training process is as follows: Entity annotation data in the ecological and environmental field is collected: pollution source names, monitoring indicator names, and geographical location names are used to construct a training dataset, which is divided into an 80% training set and a 20% test set. The BERT model uses a pre-trained bert-base-chinese model, and the LSTM model is set with a hidden layer dimension of 256 and a layer count of 2. The model is trained using the training set, and its performance is verified using the test set. Model parameters are adjusted to achieve an entity recognition accuracy of over 95%. The trained model is then deployed to a computing server for entity recognition on ecological and environmental subject data.

[0042] Based on the GeoSpark spatial big data processing framework, spatial correlation analysis is performed on the ecological and environmental data after entity identification. Different monitoring indicator data, such as PM2.5, PM10, and sulfur dioxide concentration data, are correlated at the same geographical location to generate comprehensive environmental quality data for that location. Pollution source data is also correlated with surrounding environmental monitoring data to analyze the impact of pollution sources on the surrounding environment. The correlated and fused data is stored in a data cache pool.

[0043] The data management layer includes a spatiotemporal index service module and an intelligent update module. The spatiotemporal index service module uses a distributed index architecture to connect to the data cache pool and constructs a two-dimensional and three-dimensional hybrid spatiotemporal index for the spatiotemporal characteristics of the fused ecological and environmental data. The intelligent update module is used for automatic incremental updates and historical version tracking of the ecological and environmental data.

[0044] The spatiotemporal indexing service module, based on the PostgreSQL database and the PostGIS spatial extension plugin, constructs a composite index structure of R-trees, quadtrees, and time axes. For 3D spatial data containing elevation information, such as terrain data and upper-air environmental monitoring data, an R-tree index is used. The data boundaries are defined by the MBR 3D extent calculation formula, and the 3D spatial data is hierarchically partitioned to construct the R-tree index.

[0045] For two-dimensional planar data that does not contain elevation information, such as ground environmental monitoring station data, a quadtree index is used to recursively divide the planar space into four quadrants according to longitude and latitude to index the data; the time dimension is linked to the spatial index through the time axis index, and the time attribute field of the data is added to the index to realize spatiotemporal combined query.

[0046] Meanwhile, a spatiotemporal stratification strategy is adopted to store and index data of different time periods and spatial ranges in layers, and a hot data caching mechanism is implemented. High-frequency queried data is cached in memory to optimize index query performance and ensure that the response time of spatiotemporal combination queries is controlled within 500 milliseconds.

[0047] The intelligent update module enables automatic incremental updates of data based on binlog log synchronization. When data in the data source is added, modified, or deleted, the binlog log records the data change information. The system monitors the binlog log in real time to obtain the data change content and performs incremental updates on the data cache pool and database, avoiding resource waste and update lag caused by full synchronization. Version control is used to create a version identifier for each data update, recording information such as the data update time, update content, and update operator. It supports the tracing and rollback of historical version data, making it easy for users to view the data change history.

[0048] The service publishing layer includes application service interfaces, which are connected to the spatiotemporal index service module and the intelligent update module through a load balancer. It supports scenario-based applications and the access of third-party systems through API interfaces, SDK development packages and visualization component libraries.

[0049] The application service interface is developed using Spring Boot and integrated with the Nginx load balancer to achieve request distribution and load balancing, thereby improving service availability and concurrency capabilities. It provides API interfaces, SDK development packages, and a visual component library to facilitate scenario-based applications, such as intelligent ecological environment monitoring platforms, emergency decision-making and command platforms, and third-party system access. Through the interface permission management module, access control is implemented for connected applications and systems to ensure data security.

[0050] The deployment of edge node devices involves using an embedded industrial computer equipped with an ARM Cortex-A57 quad-core processor, 16GB of RAM, and a 512GB solid-state drive. This computer is used to deploy lightweight ETL tools for edge preprocessing units, enabling localized preliminary data processing. Edge node devices are distributed at various data collection points, such as pollution source monitoring stations and environmental monitoring sensor deployment points. They are connected to the collection equipment via wired or wireless means to collect various types of ecological and environmental data.

[0051] The server cluster consists of data servers, compute servers, and application servers. The data servers employ a distributed storage architecture, configuring multiple high-performance servers to form a storage cluster with a total storage capacity of no less than 100TB. This cluster stores the collected raw data, preprocessed data, fused data, and historical versions of data. The compute servers utilize GPU-accelerated servers to run the improved BERT-LSTM hybrid model of the data governance engine module, the data quality assessment model, and the spatiotemporal index construction algorithm, providing robust computing power. The application servers deploy application service interfaces, load balancers, and other components of the service publishing layer, providing external service support.

[0052] Gigabit Ethernet switches and routers are used to build a high-speed and stable internal network, ensuring data transmission rate and stability between edge node devices and server clusters, and between servers within the server cluster; firewall devices are configured to ensure system network security.

[0053] The above are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. An ecological multi-source big data fusion system for theme and scene application, characterized in that, The data collection layer, the data fusion layer, the data management layer and the service publishing layer interact through a distributed bus; The data collection layer includes a multi-protocol adaptation unit and an edge preprocessing unit. The multi-protocol adaptation unit collects all types of ecological environment data through databases, networks, files and interfaces. The edge preprocessing unit cleans, removes duplicates, processes outliers and converts standards for ecological environment data, and outputs high-quality standardized data. The data fusion layer includes a data cache pool and a data governance engine module. The data cache pool and the data governance engine module are communicatively connected. The data governance engine module uses entity recognition and spatial correlation analysis to correlate and fuse ecological environment data under the same theme, generates ecological environment fusion data, and temporarily stores it in the data cache pool. The data management layer includes a spatio-temporal index service module and an intelligent update module. The spatio-temporal index service module connects the data cache pool using a distributed index architecture, and constructs a two-three-dimensional hybrid spatio-temporal index for the spatio-temporal characteristics of the fused ecological environment data. The intelligent update module is used for automatic incremental updating and historical version tracing of ecological environment data. The service publishing layer includes an application service interface. The application service interface connects the spatio-temporal index service module and the intelligent update module through a load balancer, supports scenario-based applications through API interfaces, SDK development kits and visual component libraries, and supports the access of third-party systems. 2.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 1, characterized in that: The multi-protocol adaptation unit includes data transmission protocols matching HTTP / HTTPS, MQTT, Kafka, FTP / SFTP and JDBC / ODBC, and reserves custom protocol extension interfaces. 3.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 1, characterized in that: The edge preprocessing unit includes a lightweight ETL tool implanted in an embedded computing chip, which performs localized preliminary processing on ecological environment data, including key field integrity verification, data validity verification, and lossless compression processing of video and image large files. 4.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 1, characterized in that: The data collection layer also includes a dynamic rule engine unit, which uses a visual drag-and-drop rule configuration interface to customize input data cleaning rules, including numerical range rules, logical verification rules and correlation verification rules.

5. The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 4, characterized in that: The data collection layer also includes a data quality assessment unit, which assesses the integrity, accuracy, consistency and timeliness of ecological environment data through a data quality assessment model. When the assessment score is below the threshold, the assessed ecological environment data is sent to the manual review process. 6.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 5, characterized in that: The comprehensive score calculation of the data quality evaluation model satisfies the following formula: ; wherein S is a data quality comprehensive score, is a score for completeness, accuracy, consistency and timeliness, is a weight corresponding to the basic dimension, ; is a score for N custom dimensions, is a weight corresponding to the custom dimension, . 7.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 1, characterized in that: The entity recognition in the data governance engine module is implemented using an improved BERT-LSTM hybrid model, which satisfies the following logic: where y is the predicted label sequence, A is the transition probability matrix, denotes the probability of transitioning from label to label , is the probability of the i-th word of the model output corresponding to label y_i, and Z(x) is the normalization factor. 8.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 1, characterized in that: The spatio-temporal index service module uses a composite index structure of R-tree, quad-tree and time axis. The three-dimensional spatial data of elevation information uses R-tree index, which divides the three-dimensional spatial data hierarchically through the minimum bounding rectangle. Two-dimensional plane data uses quad-tree index, which recursively divides the plane space into four quadrants according to longitude and latitude. The time dimension is associated with the spatial index through the time axis index. In the index construction process, a spatio-temporal hierarchical strategy and a hot data caching mechanism are used to realize dynamic optimization. The response time of spatio-temporal combined query is controlled within 500 milliseconds. 9.The theme and scenario application oriented ecological environment multi-source big data fusion system according to claim 8, characterized in that: In the R-tree index construction of the three-dimensional spatial data, the three-dimensional range of the minimum bounding rectangle (MBR) satisfies the following relationship: wherein, , are minimum and maximum coordinate values in the x-axis direction of the three-dimensional spatial data, respectively, , are minimum and maximum coordinate values in the y-axis direction, respectively, , are minimum and maximum coordinate values in the z-axis direction, respectively.