Monitoring Method and Device for Environmental Impact of Territorial Spatial Planning

Through the combination of a self-supervised space-time map mask transmission attention network and multimodal diffusion converter, the problems of multi-source heterogeneous environmental monitoring data fusion and feature extraction are solved, and the precise monitoring of environmental impact and the application of deep learning early warning mechanisms is realized, and the monitoring effect is improved.

CN119848703BActive Publication Date: 2025-06-27BEIJING ZHONGLIAN WORLD CONSTR PLANNING & DESIGN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510328302.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate multi-source heterogeneous environmental monitoring data, the feature extraction is insufficient and the early warning mechanism is incomplete, which cannot meet the needs of real-time monitoring of environmental impacts in land space planning.

Method used

The self-supervised space-time map mask transmission attention network is used to extract space-time features, combine with multimodal diffusion converters for feature fusion, and accurately identify and timely early warning of environmental risks through deep learning early warning mechanism.

Benefits of technology

The effective fusion of multi-source heterogeneous data is achieved, the accuracy of feature extraction is improved, the ability to express environmental changes is enhanced, and the monitoring effect is improved through the deep learning early warning mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848703B_ABST
    Figure CN119848703B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for monitoring the environmental impact of territorial spatial planning. The method includes: collecting and preprocessing multi-source environmental monitoring data; using a self-supervised spatio-temporal graph mask transfer attention network to extract spatio-temporal features from the preprocessed multi-source environmental monitoring data to obtain initial spatio-temporal features; performing fusion processing on the initial spatio-temporal features based on a multi-modal diffusion transformer to obtain fusion features; calculating environmental impact assessment indicators based on the fusion features; and generating dynamic monitoring results including spatio-temporal evolution diagrams, trend analysis charts, and warning information according to the environmental impact assessment indicators based on a preset visualization template and warning rules. The present application improves the accuracy and efficiency of environmental impact monitoring through deep learning and multi-modal data fusion technologies, and realizes the accurate assessment and timely warning of the environmental impact during the implementation of territorial spatial planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of environmental monitoring, and particularly to a method and device for monitoring the environmental impact of territorial space planning. Background Art

[0002] With the acceleration of urbanization and the improvement of the intensity of territorial space development and utilization, the importance of environmental impact monitoring in territorial space planning has become increasingly prominent. Environmental monitoring data has characteristics such as multi-source heterogeneity and strong spatio-temporal heterogeneity, and a systematic monitoring and evaluation system needs to be established.

[0003] Currently, the commonly used environmental monitoring methods mainly include change detection based on remote sensing images and parameter collection based on ground monitoring stations. These methods are relatively mature in the processing of single data sources, but there are still deficiencies in the fusion of multi-source data.

[0004] Existing technologies usually use simple feature splicing or weighted average methods for data fusion, and conduct environmental impact assessment through traditional statistical models. This method is difficult to fully exploit the spatio-temporal correlation features between data, and information loss is likely to occur when processing high-dimensional heterogeneous data.

[0005] In addition, existing environmental monitoring systems generally have problems such as insufficient feature extraction and imperfect early warning mechanisms, making it difficult to achieve accurate identification and timely early warning of environmental changes, and unable to meet the requirements of real-time monitoring of environmental impacts in territorial space planning. Summary of the Invention

[0006] In view of this, this application provides a method and device for monitoring the environmental impact of territorial space planning, which solves the problems in the prior art of being difficult to effectively fuse multi-source heterogeneous data, insufficient feature extraction, and imperfect early warning mechanism.

[0007] An embodiment of this application provides a method for monitoring the environmental impact of territorial space planning, including:

[0008] Collect and preprocess multi-source environmental monitoring data, where the multi-source environmental monitoring data includes at least two of satellite remote sensing images, meteorological station data, land cover data, and ecological environment monitoring data;

[0009] Use a self-supervised spatio-temporal graph mask transfer attention network to extract spatio-temporal features from the preprocessed multi-source environmental monitoring data to obtain initial spatio-temporal features;

[0010] Based on a multi-modal diffusion transformer, perform fusion processing on the initial spatio-temporal features to obtain fusion features;

[0011] Calculate environmental impact assessment indicators based on the fusion features;

[0012] Based on a preset visualization template and warning rules, dynamic monitoring results including spatio-temporal evolution maps, trend analysis charts, and warning information are generated according to the environmental impact assessment indicators.

[0013] An embodiment of this application also provides a device for monitoring the environmental impact of territorial spatial planning, including:

[0014] A data acquisition and preprocessing module, configured to acquire and preprocess at least two types of multi-source environmental monitoring data including satellite remote sensing images, meteorological station data, land cover data, and ecological environment monitoring data;

[0015] A spatio-temporal feature extraction module, configured to use a self-supervised spatio-temporal graph mask transfer attention network to extract spatio-temporal features from the preprocessed multi-source environmental monitoring data to obtain initial spatio-temporal features;

[0016] A feature fusion module, configured to perform fusion processing on the initial spatio-temporal features based on a multi-modal diffusion transformer to obtain fusion features;

[0017] An evaluation index calculation module, configured to calculate environmental impact assessment indicators including land use change index, ecosystem damage degree, and comprehensive environmental quality index based on the fusion features;

[0018] A visualization and warning module, configured to generate dynamic monitoring results including spatio-temporal evolution maps, trend analysis charts, and warning information according to the environmental impact assessment indicators based on a preset visualization template and warning rules.

[0019] This application has the following technical effects:

[0020] The effective fusion of multi-source heterogeneous data is realized through a self-supervised spatio-temporal graph mask transfer attention network, improving the accuracy of feature extraction; the multi-modal diffusion transformer is used for feature fusion processing, enhancing the expression ability of environmental change features; the warning mechanism based on deep learning realizes the accurate identification and timely warning of environmental risks, improving the monitoring effect. Description of the Drawings

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments:

[0022] Figure 1 is a schematic flowchart of the method for monitoring the environmental impact of territorial spatial planning provided by an embodiment of this application;

[0023] Figure 2 is a schematic diagram of the process of acquiring and preprocessing multi-source environmental monitoring data in an embodiment of this application;

[0024] Figure 3 is a schematic flowchart of the specific steps of step S2 in an embodiment of this application;

[0025] Figure 4 It is a schematic diagram of the specific process of step S3 in the embodiment of the present application;

[0026] Figure 5 It is a schematic diagram of the specific process of step S4 in the embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of the specific process of step S5 in the embodiment of the present application;

[0028] Figure 7 It is a schematic diagram of the environmental impact monitoring device for territorial space planning in the embodiment of the present application. Detailed implementation manners

[0029] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0030] As Figure 1 shown, the embodiment of the present application provides a method for environmental impact monitoring of territorial space planning, including:

[0031] S1: Collect and preprocess multi-source environmental monitoring data, where the multi-source environmental monitoring data includes at least two of satellite remote sensing images, meteorological station data, land cover data, and ecological environment monitoring data;

[0032] As Figure 2 shown, in the multi-source environmental data collection stage of the embodiment of the present application, it can be carried out through a pre-established data collection system, which has both data collection and data preprocessing functions (such as data cleaning, format unification, spatial registration, and time synchronization, etc.). For satellite remote sensing image data, the embodiment of the present application uses high-resolution satellite remote sensing images with a spatial resolution better than 1 meter. Such remote sensing images mainly come from high-resolution series satellites or commercial remote sensing satellites. Their high spatio-temporal resolution characteristics enable them to clearly capture subtle changes on the ground surface, which helps to detect environmental changes in a timely manner. To ensure the continuity and timeliness of monitoring, the acquisition frequency of remote sensing images is set to 1-2 times per month.

[0033] For the collection of meteorological station data, it mainly includes key meteorological parameters such as temperature, precipitation, and humidity. The embodiment of the present application preferably uses the observation data of automatic weather stations to achieve all-weather automatic collection. Considering the spatial variability of meteorological elements, a differential strategy is adopted in the layout of meteorological stations: one station is set every 50-100 square kilometers in urban areas, while in rural areas, it can be appropriately relaxed to one station every 200-300 square kilometers to ensure the spatial representativeness of data collection.

[0034] The acquisition of land cover data is mainly carried out through multi-temporal remote sensing images. Land cover information is obtained through land cover classification methods, and the classification includes main types such as forest land, grassland, cultivated land, construction land, water area, etc. To ensure the accuracy of land cover data, the embodiments of this application also verify and improve the classification results by combining field survey data.

[0035] The acquisition of ecological environment monitoring data is more comprehensive, including monitoring data of environmental elements such as air quality, water quality, and soil. Specifically, air quality monitoring mainly collects concentration data of pollutants such as PM2.5, PM10, SO2, NO x etc.; water quality monitoring includes indicators such as pH value, dissolved oxygen, total nitrogen, and total phosphorus of surface water bodies; soil monitoring focuses on key parameters such as heavy metal content and organic matter content. These data are usually obtained through fixed monitoring stations or portable monitoring devices.

[0036] In the data preprocessing stage, the embodiments of this application adopt a multi-stage processing strategy. First, geometric correction and radiometric calibration are performed on the remote sensing images. Geometric correction uses the control point registration method to ensure that the spatial position accuracy is better than 1 pixel; radiometric calibration converts the image gray value into surface reflectance through the atmospheric radiative transfer model. For meteorological and environmental monitoring data, outlier detection and missing value processing are mainly carried out. Outlier detection adopts the 3σ criterion to mark the data beyond the normal range; missing values are supplemented by time series interpolation methods to ensure the continuity of the data.

[0037] In addition, the embodiments of this application innovatively introduce the spatio-temporal data cube structure, which organizes monitoring data from different sources into a three-dimensional data body with the same spatio-temporal resolution. In the spatial dimension, regular grid division is adopted, and the grid size is determined according to the size of the study area and data characteristics, usually 100 meters × 100 meters; in the time dimension, an appropriate time interval is set according to the monitoring frequency, which can be days, weeks or months. Through this unified data organization method, the effective integration of multi-source heterogeneous data is achieved.

[0038] To ensure data quality, the embodiments of this application also establish a complete quality control system, including instrument calibration and operation and maintenance specifications in the data acquisition link, quality assessment indicators in the preprocessing link, and cross-validation mechanisms for processing results. Through these measures, the reliability and accuracy of monitoring data can be significantly improved.

[0039] Finally, for the spatio-temporal alignment problem of different types of data, the embodiments of this application adopt a fusion strategy based on spatial interpolation and time weighting. For point observation data with uneven spatial distribution (such as meteorological station data), the Kriging interpolation method is used to generate a continuous spatial distribution field; for data with inconsistent time resolutions, uniform processing is achieved through the time weighted average method. This processing method not only preserves the characteristics of the original data but also ensures the comparability and consistency between different data sources.

[0040] S2: Use the self-supervised spatio-temporal graph mask transfer attention network to extract spatio-temporal features from the preprocessed multi-source environmental monitoring data to obtain initial spatio-temporal features;

[0041] As Figure 3 shown, S2 specifically includes:

[0042] S2.1: Based on the preprocessed multi-source environmental monitoring data, divide the monitoring area into spatio-temporal unit nodes, establish spatial edges based on geographical location adjacency, establish time edges based on time series continuity, and construct a spatio-temporal relationship graph;

[0043] The self-supervised spatio-temporal graph mask transfer attention network first needs to construct a spatio-temporal relationship graph. In the spatial dimension, divide the monitoring area into regular grid cells, and each grid cell serves as a node in the graph. The initial feature vector of the node includes the depth features of the remote sensing image at that location (extracted by a pre-trained convolutional neural network) and environmental parameters. In terms of establishing spatial edges, an 8-neighborhood connection method is adopted, that is, each node is connected to its adjacent nodes in the up, down, left, right, and four diagonal directions. This connection method can effectively capture local spatial dependence relationships. In the time dimension, the corresponding nodes at different time points are connected by time edges to form a time series association. The weight of the time edge can be adjusted according to the distance of the time interval to reflect the time decay effect.

[0044] S2.2: For the nodes of the spatio-temporal relationship graph, randomly mask 15%-30% of the node features, and reconstruct the masked node features through the information of surrounding nodes for self-supervised pre-training to obtain a pre-trained attention network model;

[0045] To improve the model's ability to understand environmental change features, the embodiments of this application design a self-supervised learning strategy based on masking. Specifically, in each training iteration, 15%-30% of the nodes are randomly selected for feature masking. The setting of the masking ratio is based on experimental verification. A too low masking ratio makes it difficult for the model to learn effective feature representations, while a too high masking ratio may lead to excessive information loss. For the masked nodes, their connection relationships in the graph are retained, but their feature vectors are set to special mask tokens. The model needs to reconstruct the features of these masked nodes through the information of the unmasked nodes around them. This reconstruction task forces the model to learn the intrinsic association patterns between nodes. At the same time, a contrastive learning loss function is introduced to maximize the similarity of features at the same position at different time points and minimize the similarity of features at spatially uncorrelated positions, enhancing the model's perception ability of spatio-temporal patterns.

[0046] S2.3: Input the spatio-temporal relationship graph into the pre-trained attention network model, calculate the attention weights between spatially adjacent nodes through the spatial attention layer, and model the long-range dependence relationship between different time points through the temporal attention layer to obtain the node attention weights;

[0047] In S2.3, after obtaining the pre-trained model, a two-layer attention mechanism is used to capture complex spatio-temporal dependence relationships. The spatial attention layer calculates the attention weights based on the similarity of node features. Specifically, the scaled dot-product attention mechanism is adopted, that is, the query vector, key vector, and value vector are obtained through linear transformations respectively, then the inner product of the query vector and the key vector is calculated, and after softmax normalization, it is weighted and summed with the value vector. This enables the model to adaptively focus on nodes with strong spatial correlation. The temporal attention layer adopts a similar mechanism, but establishes connections between different time points, and can effectively capture long-term time dependence relationships, such as seasonal change patterns or interannual change trends.

[0048] The embodiments of this application take the Yangtze River Delta urban agglomeration as an example to demonstrate the application of the attention model. For example, when analyzing the environmental changes in a new area, the spatial attention layer finds that the influence weight of the Lingang New Area on the Pudong Central Area is 0.42 by calculating the correlation degree between the central point and the surrounding areas, indicating a strong environmental connection between the two areas. While the attention weight with Yangpu District located in the northwest direction is only 0.15, indicating that the spatial correlation of environmental influence is relatively weak.

[0049] In the case of Suzhou Industrial Park, the temporal attention layer identifies key correlation patterns with historical data. For example, the attention weight between the ozone pollution event in the summer of 2023 and the data in the same period of 2022 is as high as 0.68, revealing strong seasonal characteristics. While the attention weight with the data in the winter of 2021 is only 0.23, indicating a relatively low correlation of environmental changes between different seasons. This temporal attention mechanism effectively captures the long-term seasonal change patterns.

[0050] In the analysis of Nanjing Jiangbei New Area, the model simultaneously uses the multi-head attention mechanism. One attention head focuses on the diffusion pattern of industrial pollutants and finds that the attention weight with the surrounding chemical industrial parks is 0.55; another attention head focuses on the ecological system connectivity, and the attention weight with the wetlands along the Yangtze River reaches 0.61. This multi-head design enables the model to simultaneously focus on different types of environmental impact factors.

[0051] In the monitoring of the Hangzhou Bay industrial belt, the attention model reveals complex spatio-temporal dependence relationships through hierarchical attention calculation. The attention weights between the coastal industrial areas and the inland ecological protection areas show a gradient decreasing feature, gradually decreasing from 0.48 near the shore to 0.12 at the far end. This spatial attenuation feature highly coincides with the actual observed pollutant diffusion pattern. At the same time, the model detects a time lag of 3 - 5 days between the peak period of industrial activities and the air quality change, and the corresponding temporal attention weights fluctuate between 0.35 - 0.52.

[0052] These examples fully demonstrate the practical application effect of the attention model in environmental monitoring. The spatial attention layer can accurately quantify the intensity of environmental impact transmission between different regions, while the temporal attention layer successfully captures the time evolution law of environmental changes. The introduction of the multi-head attention mechanism further improves the model's understanding ability of complex environmental systems, making the environmental impact assessment results more accurate and reliable.

[0053] S2.4: Based on the node attention weights, iteratively update the node features, fuse the node features within the spatial neighborhood, and use the gated recurrent unit to process the temporal dimension feature evolution to obtain the initial spatio-temporal features.

[0054] In S2.4, during the feature update phase, the node features are iteratively updated based on the calculated attention weights. Specifically, each node aggregates the information of its neighboring nodes according to the attention weights to achieve dynamic feature updates. This process is repeated multiple times to enable sufficient information transmission in the graph. For the fusion of node features within the spatial neighborhood, a weighted summation method is adopted, and the weights are determined by the attention scores. In the time dimension, a gated recurrent unit (GRU) is used to process the evolution of features. The update gate and reset gate mechanisms of the GRU can selectively retain or update historical information, effectively modeling the dynamic changes in the time series. Finally, through multi-scale feature integration, features at different spatial ranges (such as local, regional, and global) and different time scales (such as daily, monthly, and quarterly) are combined to obtain a comprehensive spatio-temporal feature representation.

[0055] The embodiment of this application can execute step S2.4 through a complete self-supervised spatio-temporal graph masked transfer attention network structure. This network is based on a spatio-temporal relationship graph and includes three main components: a feature mask layer, a dual attention layer, and a feature update layer. The feature mask layer achieves self-supervised learning through random masking; the dual attention layer processes spatial and temporal dependencies respectively; and the feature update layer is responsible for the dynamic update and evolution of node features. The entire network is trained in an end-to-end manner to achieve efficient extraction of spatio-temporal features from the original monitoring data.

[0056] The advantages of this network structure are mainly reflected in three aspects: First, the self-supervised learning mechanism reduces the dependence on labeled data and improves the generalization ability of the model; second, the double attention structure can adaptively capture spatio-temporal dependencies at different scales; finally, the iterative feature update ensures sufficient information transmission and global consistency of feature expression. Experimental results show that this method has achieved significant performance improvement in the environmental change detection task, with the average accuracy increased by more than 15%.

[0057] Taking the environmental monitoring of the Yangtze River Delta urban agglomeration as an example, this region includes many large and medium-sized cities such as Shanghai, Suzhou, and Wuxi, and is a typical complex region where urbanization coexists with ecological environment protection.

[0058] When constructing the spatio-temporal relationship graph (S2.1), first, the 25,000-square-kilometer monitoring area is divided into 1km×1km grid cells, and each cell serves as a node in the graph. For example, the node (116, 234) located in the Suzhou Industrial Park establishes spatial edge connections with its surrounding 8 adjacent nodes, and these nodes are distributed on the park boundary and the surrounding ecological conservation areas. In the time dimension, monthly observation data from 2020 to 2024 are used to establish time edges, connecting nodes at the same location but different time points to form a continuous time chain.

[0059] In the self-supervised pre-training stage (S2.2), random masking is performed on the constructed spatio-temporal relationship graph. In each training iteration, 20% of the nodes are randomly selected for masking. For example, the environmental characteristics of a certain node in Suzhou Industrial Park in July 2023 [PM2.5 = 45 μg / m³, SO2 = 12 μg / m³, NO2 = 38 μg / m³, vegetation coverage = 0.35] are replaced with the mask token [MASK]. The model reconstructs these masked features by analyzing the information of the unmasked nodes around, such as adjacent ecological conservation area nodes and historical month data. After a large number of iterative trainings, the model gradually masters the spatial correlation and seasonal variation rules of regional environmental elements.

[0060] In the attention weight calculation stage (S2.3), through the analysis of the spatial attention layer, it is found that the attention weight of the industrial park node to the eastern residential area node is 0.12, to the western ecological conservation area node is 0.35, and to the southern Taihu Lake water area node is 0.28. This reflects the influence degree of different functional areas on the environmental quality of the park. The temporal attention layer identifies a significant long-range dependence relationship (attention weight > 0.4) between this node in summer (June - August) and winter (December - February) every year, which reveals the coupling effect of seasonal industrial activities and climate change.

[0061] In the above case of the industrial park, the calculation process of the attention weight fully reflects the model's ability to depict the environmental impact transmission mechanism. For the calculation of spatial attention, the model first constructs a 128-dimensional feature vector, which contains key information such as pollutant concentration, land use type, and meteorological parameters of the industrial park node, and adds spatial position information through sine position encoding. Through a learnable parameter matrix, the model transforms these original features into query vectors, key vectors, and value vectors respectively to prepare for subsequent attention calculation.

[0062] In the specific calculation process, the model calculates the attention scores between the industrial park node and the surrounding areas through dot product operation and scaling operation. After softmax normalization, the attention weights of 0.12 to the eastern residential area, 0.35 to the western ecological conservation area, and 0.28 to the southern Taihu Lake water area are obtained. These values intuitively reflect the influence degree of different functional areas on the environmental quality of the industrial park. Among them, the relatively high weight (0.35) of the western ecological conservation area indicates that it plays an important buffering role in the park environment through ecological regulation.

[0063] In terms of the calculation of temporal attention, the model processed the monthly monitoring data series from 2020 to 2024. Through a similar attention mechanism, the model identified seasonal environmental change patterns. Specifically, when analyzing the summer nodes in 2023, the model calculated that the attention weight between it and the data in the same period of the previous year was as high as 0.68, while the weight with the winter data of the current year was 0.42. This significant seasonal correlation reflects the periodic change characteristics of industrial activity intensity and meteorological conditions.

[0064] Finally, the model output two key attention matrices: one describes the spatial correlation intensity, and the other depicts the time dependence relationship. These attention weights not only quantify the transmission intensity of environmental impacts but also provide an interpretable analysis basis for environmental management. For example, based on these weights, the management department can prioritize attention to the boundary management of ecological conservation areas and adjust the environmental supervision strategy in summer accordingly. This analysis method based on the attention mechanism successfully realizes the accurate identification and quantitative expression of key impact relationships in complex environmental systems.

[0065] In the feature update and evolution stage (S2.4), based on the calculated attention weights, the model iteratively updates the features of the industrial park nodes. For example, considering the influence of Taihu Lake waters, humidity and temperature factors are incorporated into the feature set; considering the regulatory role of ecological conservation areas, the vegetation purification ability factor is added. Using gated recurrent units to process the time series from 2020 to 2024, it successfully captures the improvement trend that the annual average concentration of PM2.5 decreased from 52 μg / m³ in 2020 to 38 μg / m³ in 2024 and the periodic change law of the increase in summer ozone concentration with the industrial upgrading of the park and the implementation of environmental protection measures. The finally obtained spatio-temporal features not only reflect the current environmental conditions of the industrial park but also contain the dynamic process of environmental quality improvement and future trend prediction.

[0066] S3: Fuse and process the initial spatio-temporal features based on a multimodal diffusion transformer to obtain fused features;

[0067] As Figure 4 shown, S3 specifically includes:

[0068] S3.1: Perform feature encoding on the initial spatio-temporal features and add modality type tags to obtain encoded features with modality tags;

[0069] The first key step of the multi-modal diffusion transformer is feature encoding. For the initial spatio-temporal features, a multi-layer perceptron is used for non-linear transformation to map them into a 256-dimensional feature space. For remote sensing image data, Vision Transformer is used as the backbone network to divide the image into 16×16 image patches, and hierarchical visual features are extracted through positional encoding and self-attention mechanism. For meteorological data, considering its temporal characteristics, a bidirectional long short-term memory network (BiLSTM) is used for encoding to capture the temporal dependencies of meteorological elements. In addition, different source data can be distinguished through a dedicated modality type tagging mechanism, that is, unique modality identifiers are added to the feature vectors. These identifiers are in the form of learnable embedding vectors and are adaptively optimized during the training process, which helps the model identify and process features of different modalities.

[0070] In the environmental monitoring of territorial space planning, traditional methods are difficult to effectively process multi-source heterogeneous data, especially the fusion problem of different modality data such as remote sensing images and meteorological data. To address this challenge, the embodiments of this application provide an innovative multi-modal feature encoding structure. This structure adopts a targeted feature extraction network to ensure that the features of different types of data can be expressed and fused in a unified feature space.

[0071] Taking a typical monitoring area in the Yangtze River Delta region as an example, the model first processes the remote sensing image data. Through the Vision Transformer network, the 30-meter resolution Landsat satellite image is divided into 16×16 pixel image patches. This division method enables the model to simultaneously focus on local texture features and regional spatial structures. After being processed by 8 layers of Transformer encoders, hierarchical features reflecting land cover changes are successfully extracted. Experiments show that compared with traditional CNN methods, the accuracy of this processing method is increased by 23% when detecting subtle land use changes.

[0072] For meteorological data, considering its strong temporal correlation, the model uses a bidirectional LSTM network for encoding. Taking the hourly observation data of meteorological elements such as temperature, precipitation, and wind speed as an example, through a 3-layer BiLSTM structure (with a hidden layer dimension of 128), the model can consider both forward and backward temporal dependencies. This bidirectional processing mechanism enables the model to accurately capture the seasonal changes and sudden anomalies of meteorological elements, and the warning accuracy for extreme weather events reaches 85%.

[0073] To achieve the unified representation of multi-source data, the model innovatively introduces learnable modality identifiers. Specifically, [IMG] identifiers are added to remote sensing data, and [MET] identifiers are added to meteorological data. These identifiers exist in the form of 32-dimensional embedding vectors. During the training process, these identifiers are continuously optimized through backpropagation, gradually learning the feature expression rules of different data modalities. Experiments have proven that this identifier mechanism significantly improves the model's processing ability for data from different sources, and the feature fusion effect has been improved by 31%.

[0074] Finally, all features are uniformly mapped to a 256-dimensional feature space through a multi-layer perceptron. In this space, the data features of different modalities are effectively aligned, laying a foundation for subsequent environmental impact assessment. Through practical application verification in Suzhou Industrial Park, this multi-modal feature encoding scheme has successfully improved the overall accuracy of environmental monitoring to 92%, a 25 percentage point increase compared to traditional methods.

[0075] S3.2: Construct a multi-layer diffusion transformer including a diffusion attention layer and a cross-modal interaction block, input the encoded features with modality tags into the multi-layer diffusion transformer, and obtain initial transformed features;

[0076] The core of the multi-layer diffusion transformer lies in its unique dual attention mechanism. The diffusion attention layer is responsible for the smooth transfer of features, and its basic principle is to simulate the physical diffusion process. Specifically, features are regarded as particles in a diffusion medium, and the diffusion coefficient matrix is defined to control the transfer rate of features between different positions. The calculation of the diffusion coefficient takes into account two factors: spatial distance and feature similarity, ensuring stronger information exchange between similar features. The cross-modal interaction block adopts a multi-head cross-attention mechanism, allowing features of different modalities to directly perform information interaction. Each attention head focuses on different aspects of the features, and comprehensive feature representations can be obtained by integrating the outputs of multiple attention heads. This structural design effectively solves the problems of strong heterogeneity and uneven information distribution of multi-modal data.

[0077] S3.3: Gradually add Gaussian noise to the initial transformed features through forward diffusion, and then achieve feature reconstruction through reverse diffusion to obtain reconstructed features;

[0078] In S3.3, the conditional diffusion process is one of the innovation points of this application embodiment. It applies the idea of traditional diffusion models to the feature fusion task. First, features of different modalities are mapped to a unified 512-dimensional latent space through a projection layer. During the forward diffusion process, Gaussian noise is gradually added to the features, and the variance of the noise increases according to a predefined schedule. This process can be expressed as:

[0079]

[0080] where is the noise coefficient at time step t, and ε is the standard Gaussian noise. The reverse diffusion process restores the contaminated features step by step by training a denoising network. The conditional input of the denoising network includes the time step information and the modality type label, enabling it to perform adaptive denoising according to the characteristics of different modalities. This diffusion-based feature fusion method has good uncertainty modeling ability and can generate more robust fused features.

[0081] S3.4: Perform temporal alignment and spatial alignment on the reconstructed features to obtain registered features, calculate the modality reliability weights of the registered features, and perform feature dynamic fusion based on the modality reliability weights to obtain the fused features.

[0082] In S3.4, the feature registration step adopts a dual alignment strategy. Temporal alignment uses the dynamic time warping (DTW) algorithm to handle the problem of inconsistent sampling frequencies of different modality data. Spatial alignment is based on the spatial transformation network (STN) to learn the spatial transformation relationship between different modality features. The calculation of the reliability weights considers three aspects: signal-to-noise ratio, data integrity, and spatio-temporal consistency. Specifically, the signal-to-noise ratio is evaluated by the ratio of the variance of the feature to the noise level; data integrity is calculated based on the proportion of missing values; and spatio-temporal consistency is measured by the gradient continuity of the feature in the spatio-temporal dimension. The final fused feature is obtained through weighted average, and the dynamic adjustment of the weights ensures that more reliable modalities have a greater contribution in the fusion process.

[0083] In the environmental monitoring practice in the Yangtze River Delta region, the sampling frequencies of different data sources vary significantly: satellite remote sensing data is usually once every 15 - 30 days, air quality monitoring is once per hour, land cover data may be updated quarterly, and meteorological observation data is once every 10 minutes. This inconsistency in time scales poses challenges to data fusion. To address this issue, this application improves the traditional DTW algorithm to make it more suitable for the environmental monitoring scenario.

[0084] Specifically, the improved DTW algorithm introduces an environmental feature weight factor. Taking the monitoring data in Suzhou Industrial Park as an example, when calculating the time series distance, the algorithm assigns different weights to different environmental indicators. For example, when calculating the alignment of the PM2.5 concentration series, the weights of meteorological conditions (such as wind speed, humidity) are set to 0.6, while the weight of land use change is set to 0.3, which reflects the influence degree of different factors on pollutant concentration. Experiments show that this weighted strategy improves the accuracy of temporal alignment by 28%.

[0085] Meanwhile, the algorithm also integrates spatio-temporal autocorrelation features. When processing the time series of a monitoring point, not only the historical data of this point is considered, but also the influence of spatially adjacent points. For example, when aligning the ozone pollution data in the industrial park in the summer of 2023, the algorithm simultaneously refers to the data of monitoring stations within a range of 5 kilometers around, and weights the contributions of adjacent stations through a spatial attenuation function (exp(-d / 5), where d is the distance). This spatio-temporal joint alignment mechanism enables the model to more accurately depict the spatio-temporal evolution law of pollutants.

[0086] In addition, the embodiment of the present application innovatively introduces an adaptive window mechanism into the DTW framework. The window size is dynamically adjusted according to the changing characteristics of environmental factors. For periods with drastic changes (such as occasional environmental pollution), a smaller alignment window (2 - 3 days) is adopted, while for stable periods, a larger window (7 - 10 days) is used. This adaptive strategy significantly improves the algorithm's response ability to environmental mutations, advancing the detection of abnormal events by an average of 4.5 days.

[0087] Through these improvements, the DTW algorithm of the present application achieves precise temporal alignment of multi-source environmental data, providing a reliable data basis for subsequent feature fusion and environmental assessment. In the practical application in the Yangtze River Delta region, the improved algorithm controls the temporal alignment error within 5%, improving the performance by 35% compared with the traditional DTW method, fully demonstrating the innovation and practical value of this solution.

[0088] The advantages of this multi-modal diffusion transformer are mainly reflected in three aspects: First, a smooth transition in the feature space is achieved through the conditional diffusion process, avoiding information loss that may occur in direct feature fusion; second, the dual attention mechanism ensures the effective extraction and fusion of local and global features; finally, the dynamic weight allocation mechanism based on reliability improves the robustness of the fusion result. Experiments show that this method performs excellently in processing high-dimensional heterogeneous data, especially maintaining stable performance in the face of complex situations such as noise interference and data loss. In the actual application test, this fusion method improves the performance index by about 20% compared with the traditional method.

[0089] Taking the ecological environment monitoring of a coastal city as an example, this area includes the coastal zone, urban built-up area, and inland ecological protection area, and it is necessary to fuse multi-source remote sensing data and ground monitoring data for environmental assessment.

[0090] In the feature encoding stage (S3.1), the embodiments of the present application first processed three main data sources: high-resolution satellite remote sensing images (2m resolution), data from 50 ground air quality monitoring stations, and data from 30 ecological monitoring sample points. For the remote sensing images, Vision Transformer was used to segment the images into image patches of 32×32 pixels, and 256-dimensional visual feature vectors containing information such as land cover and vegetation indices were extracted. For the air quality data, a BiLSTM network was used to encode the time series of six indicators such as PM2.5, SO2, and NO2 into 128-dimensional feature vectors. For the ecological monitoring data, indicators such as species diversity index and soil quality were encoded into 64-dimensional feature vectors through a multi-layer perceptron. To distinguish these three types of data, modality markers of [IMG], [AIR], and [ECO] were added respectively to form labeled encoded features.

[0091] In the stage of constructing and applying the diffusion transformer (S3.2), the model contains 6 layers of diffusion attention layers and 4 cross-modal interaction blocks. For example, when analyzing the mangrove area in the coastal zone, the diffusion attention layer found a significant correlation between the vegetation change characteristics in the remote sensing image (the normalized vegetation index decreased from 0.75 to 0.65) and the decline in species diversity at the ground ecological monitoring points (the density of a certain type of crab decreased by 30%). The cross-modal interaction block then identified the association between this change and the increase in the concentration of sea salt aerosol detected at the nearby air quality station (the annual average value increased by 20%), revealing the comprehensive performance of the coastal ecosystem affected by sea level rise.

[0092] In the diffusion process stage (S3.3), the model gradually adds Gaussian noise to the initial features. Taking the mangrove area as an example, in the forward diffusion, the original feature matrix [vegetation index = 0.65, species diversity = 0.58, aerosol concentration = 45μg / m³] was applied with increasing noise through 500 time steps. The noise variance starts from 0.0001 and increases at a rate of β_t = 0.02. In the reverse diffusion process, the trained denoising network successfully reconstructed the spatio-temporal correlation pattern of the features with an accuracy of 92%, revealing the vulnerability characteristics of the coastal ecosystem.

[0093] In the feature registration and fusion stage (S3.4), the embodiments of the present application first perform temporal alignment to unify the monitoring data of different frequencies (remote sensing data once every 15 days, air quality data once every hour, ecological monitoring once every quarter) to the monthly scale. For spatial alignment, a spatial transformation network is used to resample all data to a unified grid of 100m×100m. When calculating the modal reliability weights, it is found that the weight of remote sensing data is higher (0.5) under sunny weather, while in rainy weather, ground monitoring data is preferentially used (weight 0.6). The final fused features not only reflect the overall health status of the mangrove area but also include the impacts of multiple stress factors such as tidal activities and urban expansion.

[0094] This complete example demonstrates how the multi-modal diffusion transformer can fuse scattered environmental monitoring data into systematic environmental assessment information through multi-step processing. This method shows strong adaptability and accuracy in practical applications. Especially when dealing with complex coastal ecological environment problems, it can provide comprehensive decision-making support information. For example, based on the fusion results, the local environmental protection department successfully identified the early signs of mangrove degradation and timely adjusted the coastal development plan, implementing targeted protection measures for the mangroves, which significantly improved the stability of the regional ecosystem.

[0095] S4: Calculate environmental impact assessment indicators based on the fused features;

[0096] As Figure 5 shown, S4 specifically includes:

[0097] S4.1: Construct an attention feature extraction network based on the fused features, and extract the environmental element change features through the spatial attention mechanism and the temporal attention mechanism to obtain the environmental element change sequence;

[0098] The design of the attention feature extraction network fully considers the complexity of environmental element changes. This network adopts a multi-branch structure, including a spatial branch and a temporal branch. The spatial attention mechanism uses a non-local attention module, which can capture the dependency relationships between any two spatial positions. Specifically, when implementing, the feature map is first segmented into regular grids, and each grid cell serves as a basic computing unit. By calculating the similarity matrix of features at different positions, a spatial attention map is generated, which reflects the degree of mutual influence between different regions. The temporal attention mechanism adopts a multi-head self-attention structure, and each attention head is responsible for capturing the change patterns at different time scales. By setting receptive fields of different sizes, short-term fluctuations and long-term trends can be simultaneously concerned. Finally, the outputs of the spatial and temporal branches are fused through an adaptive gating mechanism to obtain the complete environmental element change sequence.

[0099] S4.2: Perform spatio-temporal autocorrelation analysis and change trend analysis on the environmental factor change sequence, calculate the environmental pressure index and ecological response index through a multi-layer perceptron, and obtain the initial environmental impact assessment result;

[0100] Spatio-temporal autocorrelation analysis and change trend analysis are key links in assessing environmental impacts. The local spatial autocorrelation operator is designed based on the Moran's I index, and the spatial aggregation characteristics of each location are calculated through a sliding window. The window size is dynamically adjusted according to the characteristics of the study area, usually set to range from 3×3 to 7×7. The global temporal autocorrelation operator adopts an improved LSTM structure, introducing an attention mechanism to enhance the ability to identify key time points. This method of extracting spatio-temporal correlation features can effectively identify hot spots and critical periods of environmental changes. Subsequently, a non-linear mapping is performed through a three-layer perceptron network, with 256 neurons in the first layer, 128 neurons in the second layer, and 64 neurons in the third layer, and the activation function is LeakyReLU. This progressive feature extraction structure can gradually refine the key information of environmental pressure and ecological response.

[0101] This network first constructs a spatial adjacency matrix and calculates spatial weights based on geographical distance and attribute similarity. On this basis, local spatial features are extracted through a graph convolutional network, and at the same time, the overall spatial pattern is obtained using a global pooling operation. The extraction of time evolution features is achieved through a causal convolutional network, where dilated convolutions are used to expand the receptive field to ensure the ability to perceive long-term change trends. The generation of the pressure response distribution map adopts a multi-scale feature fusion strategy, dynamically combining features at different spatial and temporal scales through an attention mechanism.

[0102] In the adaptive weight fusion algorithm of S4.2, multi-scale wavelet transform is used to decompose the pressure response distribution map. The Daubechies wavelet is selected as the basis function for 5-level decomposition to obtain feature representations of different frequency components. The calculation of feature weights considers three aspects: the significance, stability, and spatio-temporal consistency of the features. Significance is evaluated through the statistical distribution of feature activation values, stability is calculated based on the degree of fluctuation of the time series, and spatio-temporal consistency is measured through gradient continuity. The weight calculation uses softmax normalization to ensure that the sum of the contribution degrees of features at different scales is 1.

[0103] Among them, S4.2 specifically includes:

[0104] S4.2.1: Input the environmental factor change sequence into the spatio-temporal autocorrelation network, calculate the spatial aggregation characteristics and time evolution characteristics of environmental factors through the local spatial autocorrelation operator and the global temporal autocorrelation operator, and obtain spatio-temporal correlation features;

[0105] The core of the spatio-temporal autocorrelation network is the co-design of the local spatial autocorrelation operator and the global temporal autocorrelation operator. The local spatial autocorrelation operator is constructed based on the improved Getis-Ord Gi* statistic, and captures the spatial aggregation characteristics through adaptive neighborhood definition. In specific implementation, first construct the spatial weight matrix W, and its element wij is calculated based on the spatial distance and attribute similarity. This adaptive weight design can dynamically adjust the strength of the spatial dependence relationship according to the data distribution. For each location, calculate its local G statistic to form a hotspot analysis map, reflecting the spatial aggregation pattern of environmental elements.

[0106] The global temporal autocorrelation operator adopts the multi-scale temporal convolutional network (MS-TCN) structure. By setting dilated convolutional layers with different receptive field sizes, it captures the temporal evolution characteristics in the short term, medium term, and long term simultaneously. Specifically, three parallel temporal convolution branches are designed, with receptive field sizes of 3, 7, and 15 time steps respectively, and dilated convolutions with dilation rates of 1, 2, and 4 are used for feature extraction. This multi-scale structure can simultaneously focus on the change patterns at different time scales, such as seasonal fluctuations and interannual variations. Finally, through the attention mechanism, the features at different scales are dynamically fused to obtain the comprehensive temporal evolution characteristics.

[0107] S4.2.2: Non-linearly map and extract features from the spatio-temporal correlation features through a multi-layer perceptron, calculate the spatio-temporal distributions of the environmental pressure index and the ecological response index, and obtain the pressure-response distribution map;

[0108] The multi-layer perceptron is designed with a deep residual structure, including multiple residual blocks. Each residual block contains two fully connected layers, with a batch normalization layer and a ReLU activation function added in the middle, and the gradient vanishing problem is alleviated through skip connections. The network structure is 512→256→128→64 dimensions from input to output, extracting more abstract feature representations layer by layer. During the feature mapping process, a dropout mechanism (with a ratio of 0.3) is introduced to prevent overfitting. The calculation of the environmental pressure index considers factors such as land use intensity, pollutant emissions, and human activity interference, while the ecological response index comprehensively considers indicators such as vegetation cover change, biodiversity, and the ecological service function of this application embodiment. Both indices are normalized to the [0,1] interval through the Sigmoid function for subsequent analysis.

[0109] S4.2.3: Based on the pressure-response distribution map, calculate the initial environmental impact assessment result through the adaptive weight fusion algorithm.

[0110] The adaptive weight fusion algorithm adopts an attention mechanism and a dynamic weight adjustment strategy. First, multi-scale decomposition is performed on the pressure response distribution map, and wavelet transform is used to extract different frequency components. The Haar wavelet is selected as the basis function for 4-level decomposition to obtain multiple scale coefficients. Then, an attention module is constructed to calculate the importance weights of each scale feature. The attention module contains a two-layer feedforward neural network, with the input being the feature vector and the output being the weight coefficient. The weight calculation formula is:

[0111]

[0112] where W1 and W2 are learnable parameter matrices.

[0113] In the weight fusion stage, the spatio-temporal correlation and uncertainty of features are considered. The spatio-temporal correlation is quantified by calculating the autocorrelation coefficient of features in the time and space dimensions, and the uncertainty is estimated by the Monte Carlo dropout method. The final initial evaluation result of the environmental impact is obtained through weighted combination, and the weight coefficient is automatically optimized by backpropagation. The experimental results show that this adaptive fusion method improves the evaluation accuracy by about 18% compared with the fixed weight method.

[0114] To improve the interpretability of the evaluation results, a feature importance analysis module is also designed. The contribution degree of each environmental factor to the final evaluation result is calculated through SHAP (SHapley Additive exPlanations) values to help understand the influence degree of different factors. This interpretive analysis not only improves the credibility of the evaluation results but also provides an important reference for subsequent environmental management decisions. Experimental verification shows that this evaluation method exhibits good stability and accuracy on different regional and temporal scales, and the average error of the evaluation results is controlled within 10%.

[0115] Among them, S4.2.3 specifically includes:

[0116] A1. Perform multi-scale feature decomposition on the pressure response distribution map, and extract environmental pressure features and ecological response features at different scales through wavelet transform to obtain a multi-scale feature set;

[0117] The multi-scale feature decomposition adopts a multi-layer analysis framework based on the undecimated wavelet transform (UDWT). The Daubechies-4 wavelet is selected as the basis function, and five-layer wavelet decomposition is performed to obtain sub-features in different frequency bands. Each layer of decomposition can obtain an approximation coefficient and a detail coefficient, corresponding to low-frequency and high-frequency information respectively. For the environmental pressure feature, more attention is paid to the detail coefficients with higher frequencies, which can reflect the environmental pressure changes caused by short-term human activities; for the ecological response feature, more attention is paid to the low-frequency approximation coefficients, because the response of the ecosystem usually has a certain lag and cumulativeness. To maintain the integrity of the features, the original resolution is retained after each layer of decomposition, avoiding the downsampling operation in the traditional discrete wavelet transform.

[0118] Specifically, let the original pressure response distribution map be X(t), and the wavelet decomposition at the j-th layer can be expressed as: X(t) = Aj(t) + ΣDj(t), where Aj(t) is the approximation coefficient at the j-th layer and Dj(t) is the detail coefficient. By adjusting the parameters of the wavelet basis function and the number of decomposition layers, the most suitable multi-scale representation for the environmental change features can be obtained. At the same time, a threshold denoising mechanism is introduced, and the soft threshold method is used to process the wavelet coefficients to improve the signal-to-noise ratio of the features. The final multi-scale feature set contains different spatio-temporal scale information of the environmental change.

[0119] A2. Construct an adaptive weight network based on the multi-scale feature set, calculate the weight coefficients of different scale features through the attention mechanism, and obtain the feature weight matrix;

[0120] The adaptive weight network is designed with a multi-head attention mechanism. First, the features at each scale are obtained with query vectors (Query), key vectors (Key), and value vectors (Value) through three independent linear transformations. The attention calculation formula is: Attention(Q, K, V) = softmax(QK T / d)V, where d is the feature dimension. To enhance the expression ability of the model, eight attention heads are designed, and each head independently learns the feature correlations in different aspects. This multi-head design enables the model to simultaneously focus on multiple interrelationships between different scale features.

[0121] The calculation of the weight coefficients also considers the reliability and representativeness of the features. The reliability is evaluated through the signal-to-noise ratio and spatio-temporal consistency of the features, and the representativeness is measured through the significance and uniqueness of the features. Specifically, the signal-to-noise ratio is calculated by the ratio of the feature value to the background noise; the spatio-temporal consistency is measured by the gradient continuity of the features in the spatio-temporal dimensions; the significance is evaluated through the statistical distribution of the feature activation values; and the uniqueness is calculated by the mutual information between the features and other scale features. The final weight matrix is obtained through the weighted combination of these factors, ensuring the scientificity and reliability of the feature fusion.

[0122] A3. Perform weighted fusion on the multi-scale feature set and the feature weight matrix, and calculate the initial environmental impact assessment result through a deep neural network.

[0123] Feature fusion adopts a hierarchical strategy. First, perform weighted summation on the multi-scale features based on the feature weight matrix:

[0124] F = Σw i F i

[0125] where w i is the weight of the i-th scale feature, and F i is the corresponding feature representation. To capture the non-linear relationship between features, a deep residual network is designed for feature fusion. The network structure contains multiple residual blocks, each of which consists of two convolutional layers, with a batch normalization layer and a ReLU activation function added in the middle. The use of residual connections effectively alleviates the problem of gradient disappearance in deep networks.

[0126] The specific structure of the network is: input layer (the number of channels is equal to the feature dimension) → residual block 1 (256 filters) → residual block 2 (128 filters) → residual block 3 (64 filters) → fully connected layer → output layer. During the training process, the AdamW optimizer is used, the initial learning rate is set to 0.001, and the cosine annealing strategy is used for dynamic adjustment. The loss function is selected as a combination of mean squared error loss and KL divergence loss, which not only considers the error between the predicted value and the true value but also pays attention to the difference between the predicted distribution and the true distribution.

[0127] To improve the generalization ability of the model, various regularization techniques are introduced during the training process. In addition to the commonly used L2 regularization, feature dropout (ratio of 0.2) and path dropout (ratio of 0.1) are also adopted. Feature dropout randomly discards some features, forcing the model to learn more robust feature representations; path dropout randomly discards some residual connections, enhancing the ensemble learning effect of the model. Experimental results show that this multi-level fusion strategy can effectively integrate environmental change features at different scales, and the accuracy of the final evaluation result reaches 88%, which is about 23% higher than that of the single-scale method in terms of performance.

[0128] S4.3: Calculate the environmental impact assessment indicators through adaptive threshold segmentation and multi-scale feature fusion of the initial environmental impact assessment result, where the environmental impact assessment indicators include environmental stress degree, ecological fragility degree, and environmental quality index.

[0129] In S4.3, a hierarchical progressive strategy is adopted for calculating the environmental impact assessment indicators. The adaptive threshold network determines the grading criteria for each indicator through iterative optimization. The cross-validation method is used in the optimization process to avoid overfitting. Specifically, the sample data is divided into a training set and a validation set, and the threshold parameters are adjusted by minimizing the classification error. The multi-scale feature extraction uses a dilated convolutional network to obtain multi-scale receptive fields by setting different dilation rates. When fusing features, an attention mechanism is used to adaptively weight different-scale features, and the weight coefficients are automatically learned through backpropagation. The spatial clustering uses an improved DBSCAN algorithm, which can adaptively identify environmental impact areas with irregular shapes.

[0130] When generating the final assessment indicators, the environmental stress degree mainly considers the direct pressure of human activities on the environment, including factors such as land use intensity and pollutant emissions; the ecological fragility reflects the sensitivity of the ecosystem to external disturbances and is comprehensively evaluated through indicators such as vegetation coverage and biodiversity; the environmental quality index is a comprehensive measure of the overall environmental condition of the region, integrating the monitoring data of multiple environmental elements. These three indicators are normalized to the range of 0-1 for easy comparison and analysis. The experimental results show that the accuracy of this assessment method reaches more than 85%, improving the performance by about 25% compared with the traditional method.

[0131] Among them, S4.3 specifically includes:

[0132] S4.3.1: Construct an adaptive threshold network based on the initial environmental impact assessment results, and determine the grading thresholds of the environmental stress degree, ecological fragility, and environmental quality index through multi-layer iterative optimization to obtain the assessment indicator thresholds;

[0133] The adaptive threshold network adopts a deep neural network structure and determines the grading thresholds of the assessment indicators through multi-layer iterative optimization. The input of the network is the initial environmental impact assessment results, which contain data in three dimensions: environmental stress degree, ecological fragility, and environmental quality index. First, each indicator is standardized and normalized through data preprocessing to eliminate the dimensional difference. Then, a multi-layer perceptron is used to construct the basic network structure, which includes three hidden layers with the number of neurons being 256, 128, and 64 respectively, and the ReLU activation function is used. To improve the accuracy of threshold division, a residual connection mechanism is introduced to alleviate the problem of gradient disappearance.

[0134] The threshold optimization process adopts an iterative strategy, and each iteration includes two steps: forward propagation and backward update. In forward propagation, the network generates candidate thresholds according to the current parameters; in backward update, the network parameters are optimized by minimizing the objective function. The design of the objective function takes into account two objectives: minimizing the within-class difference and maximizing the between-class difference: L = λ1L within +λ2L between, where λ1 and λ2 are balance factors. To improve the stability of the threshold, a temporal smoothing constraint is also introduced to avoid drastic fluctuations in the threshold. Finally, the optimal combination of threshold parameters is determined through cross-validation.

[0135] S4.3.2: Grade the initial environmental impact assessment results according to the assessment index threshold, and extract multi-scale spatial features through a deep convolutional network to obtain a graded feature map;

[0136] The graded feature extraction adopts an improved ResNet structure. The network contains five residual blocks, and each residual block contains two 3×3 convolutional layers with the number of channels being 64, 128, 256, 512, and 1024 in sequence. To capture spatial features at different scales, a spatial pyramid pooling module (SPP) is added after each residual block, using three pooling scales of 1×1, 2×2, and 4×4. This multi-scale feature extraction strategy can simultaneously focus on local details and global structures. During the feature extraction process, dilated convolutions are used to expand the receptive field, and the dilation rates are set to 1, 2, and 4 respectively to obtain a larger range of spatial context information.

[0137] The grading process adopts a soft classification strategy, that is, calculating the probability distribution belonging to different grades for each sample instead of directly giving a hard classification result. This soft classification method can better handle the situation with fuzzy boundaries. Specifically, the softmax function is used to calculate the probability distribution. To improve the robustness of classification, label smoothing technology is introduced to avoid overconfident predictions of the model.

[0138] S4.3.3: Perform feature fusion and spatial clustering on the graded feature map, and calculate the environmental impact assessment index.

[0139] The feature fusion adopts an attention-enhanced multi-scale fusion strategy. First, double-weight the feature maps at different scales with channel attention and spatial attention. The channel attention is realized through the squeeze-and-excitation module to capture the dependencies between channels; the spatial attention is realized through the non-local block to model the long-range dependencies between spatial positions. The fused features are compressed in channels through 1×1 convolutions to obtain a feature representation with a unified dimension.

[0140] The spatial clustering adopts an improved version of the density peak clustering algorithm. First, calculate the local density ρi of each point and the minimum distance δi to the high-density points. Then, identify the density peak points as the clustering centers, and automatically determine the number of clusters through a decision graph (ρ-δ graph). To improve the stability of clustering, a spatial constraint term is introduced to ensure the spatial continuity of the clustering results.

[0141] In the application scenario of Suzhou Industrial Park, the input of feature fusion includes feature maps at three key scales. Taking a monitoring area of 5 km × 5 km as an example, these features come from the spatial resolution levels of 250 m, 500 m, and 1000 m respectively. Each scale of feature map contains 64 channels, and each channel corresponds to different environmental monitoring indicators and derived features.

[0142] In the channel attention calculation, the squeeze operation first performs global average pooling on the feature map of each channel to obtain a 64-dimensional channel descriptor. The excitation operation then calculates the channel weights through a two-layer fully connected network (with the dimensionality change of 64→16→64). Taking the PM2.5 channel as an example, its weight reflects the correlation strength with other pollutant indicators. Usually, during heavy pollution weather, the weight of this channel will automatically increase to more than 0.3. For density peak clustering, the core input parameters of the algorithm are the local density ρi and the distance δi.

[0143] The calculation of the final environmental impact assessment indicators comprehensively considers the classification results and spatial clustering information. The environmental stress degree indicator reflects the degree of pressure of human activities on the environment and is calculated by weighted average. The ecological vulnerability indicator measures the sensitivity and recovery ability of the ecosystem, considering factors such as vegetation cover and biodiversity. The environmental quality index is a comprehensive evaluation of the regional environmental conditions, integrating the monitoring data of multiple environmental elements.

[0144] Experimental verification shows that this assessment method has high accuracy and stability. By comparing with the field survey data, the overall accuracy rate of the assessment results reaches 87%, and the spatial consistency coefficient reaches 0.83. Tests on different regional and temporal scales also show good adaptability. An important advantage of this method is that it can adaptively adjust the assessment criteria to adapt to the environmental characteristics of different regions, providing a reliable scientific basis for environmental management decisions.

[0145] S5: Based on a preset visualization template and warning rules, generate dynamic monitoring results including spatio-temporal evolution maps, trend analysis charts, and warning information according to the environmental impact assessment indicators.

[0146] Among them, as Figure 6 shown, S5 specifically includes:

[0147] S5.1: Input the environmental impact assessment indicators into a dynamic visualization network, and generate an environmental impact spatio-temporal evolution map through a spatio-temporal interpolation algorithm and a gradient color mapping to obtain a visualization basic layer;

[0148] In S5.1, the dynamic visualization network is constructed using WebGL technology to achieve efficient rendering of large-scale spatio-temporal data. The spatio-temporal interpolation algorithm combines Kriging spatial interpolation and spline time interpolation methods. Ordinary Kriging method is used for spatial interpolation, and its variogram model is optimized and selected through cross-validation. Commonly used models include spherical model, exponential model, and Gaussian model. The expression of the variogram is:

[0149]

[0150] where C0 is the nugget effect, C is the sill, and a is the range. Cubic spline function is used for time interpolation to ensure the smoothness and continuity of the time series.

[0151] The design of the gradient color mapping adopts a perceptually uniform color space to ensure the scientific nature of visual expression. For the degree of environmental stress, the red series (#FF0000 to #FFCCCC) is used, for the ecological vulnerability, the green series (#006400 to #98FB98) is used, and for the environmental quality index, the blue series (#000080 to #87CEEB) is used. The design of the color band takes into account the recognizability of people with color vision defects and ensures consistency on different display devices. Layer rendering uses WebGL shader programs to achieve smooth dynamic effects through GPU acceleration. The base layer also contains background information such as terrain, water system, and administrative boundaries, which is implemented through SVG vector graphics and supports lossless scaling.

[0152] S5.2: Conduct multi-dimensional feature analysis on the visualization base layer, generate environmental change trend charts through a trend prediction model, and obtain trend analysis results;

[0153] The multi-dimensional feature analysis adopts a method combining principal component analysis (PCA) and independent component analysis (ICA). First, dimensionality reduction is performed through PCA, and the principal components with an explained variance ratio reaching 95% are retained; then independent features are extracted through ICA to identify the key driving factors of environmental change. Feature importance is evaluated through the random forest method, and the average impurity reduction of each feature is calculated. For time series data, wavelet transform is used for time-frequency analysis to identify periodic patterns and mutation points.

[0154] The trend prediction model adopts a hybrid method combining deep learning and statistical modeling. The deep learning part uses a Transformer structure, which contains 8 attention heads, and both the encoder and decoder have 6 layers. Positional encoding uses sine and cosine functions:

[0155]

[0156] The statistical modeling part uses the SARIMA model to handle the seasonal and periodic characteristics of time series. The prediction results of the two models are fused by an ensemble learning method, and the weights are dynamically adjusted according to the historical prediction accuracy of each model.

[0157] S5.3: Construct an early warning evaluation model based on the trend analysis results, identify the environmental risk level and spatial distribution through a deep learning algorithm, and generate early warning information.

[0158] The early warning evaluation model is constructed based on a deep neural network. The main body of the network adopts the ResNeXt structure, which contains multiple parallel transformation branches, and each branch uses different convolutional kernel sizes to capture features at different scales. The basic building block of the network contains a 1×1 convolution for dimensionality reduction, a 3×3 grouped convolution for feature extraction, and a 1×1 convolution for dimensionality increase, in the form of:

[0159] Y = X + Σ(T(X))

[0160] where T represents the transformation function. To improve the generalization ability of the model, dropout (ratio 0.3) and batch normalization techniques are adopted.

[0161] The identification of the environmental risk level adopts a multi-label classification method, and simultaneously predicts the risk level and the probability of spatial distribution. The risk level is divided into four levels: low risk (blue early warning), medium risk (yellow early warning), high risk (orange early warning), and extremely high risk (red early warning). The classification loss function adopts focal loss to handle the class imbalance problem: FL(pt) = -α(1 - pt)γlog(pt), where α is the class weight and γ is the focusing parameter. The prediction of the spatial distribution is realized through a semantic segmentation network, adopting the U-Net++ architecture, which has skip connections and a deep supervision mechanism.

[0162] The generation of early warning information includes multiple levels: the first level is the basic early warning information, including the risk level, the affected range, and the duration; the second level is the detailed analysis information, including the risk evolution trend, the key influencing factors, and the uncertainty analysis; the third level is the decision-making reference information, including prevention and control measures and emergency plans. The early warning information is pushed through a Web service interface, supporting multi-terminal access and real-time update. The embodiment of this application also has an automatic alarm function. When the risk exceeds the set threshold, the push of the early warning information is automatically triggered.

[0163] Experimental verification shows that this early warning method has high accuracy and practicability. Applications in multiple pilot areas show that the early warning accuracy rate reaches 86%, and the average early warning lead time is 4 days, which can provide timely and effective decision-making support for environmental risk management. The visualization effect of the embodiments of this application is intuitive and clear, and the early warning information is transmitted timely and accurately, which has been recognized by the actual application departments. Continuous system optimization and model update ensure the stable improvement of the early warning performance and provide reliable technical support for environmental risk prevention and control.

[0164] Among them, S5.3 specifically includes:

[0165] S5.3.1: Extract the time series features of the trend analysis results, and predict the evolution trend of environmental risks through a recurrent neural network to obtain a risk prediction sequence;

[0166] The time series feature extraction adopts a hybrid architecture of bidirectional LSTM (BiLSTM) and attention-enhanced Transformer. The BiLSTM network contains three layers, with 128 hidden units in each layer, which can consider the information of historical and future time steps simultaneously and capture the bidirectional time dependence of environmental risks. To handle the long-term dependence problem, a residual connection and layer normalization mechanism are introduced into the LSTM unit. The input features are first mapped to a 256-dimensional latent space through a feature embedding layer, and then the time series position information is added through position encoding. To improve the ability to identify abnormal patterns, an anomaly detection module based on variational autoencoder (VAE) is designed, which can learn the normal distribution pattern of the data and thus identify potential risk events.

[0167] In the field of environmental risk monitoring, traditional time series analysis methods often have difficulty effectively capturing complex long-term dependence relationships and abnormal patterns. To address this problem, this application designs an innovative time series feature extraction architecture, which significantly improves the early warning ability of environmental risks through the synergistic effect of BiLSTM and Transformer.

[0168] In the actual application in Suzhou Industrial Park, the model first preprocesses and embeds the features of the original environmental monitoring data. Taking the monitoring data in 2023 as an example, the input includes the concentrations of 12 pollutants collected hourly, 8 types of meteorological parameters, and 4 types of land use change indicators. Through the feature embedding layer, these original features are mapped to a 256-dimensional latent space. Experiments show that this high-dimensional representation can better capture the interactions between environmental elements, and the feature expression ability is improved by 42%.

[0169] In the design of the BiLSTM network, a residual connection mechanism was innovatively introduced. Specifically, a direct cross-layer connection was added between LSTM units with 128 dimensions in each layer. This design significantly alleviated the vanishing gradient problem, enabling the model to effectively process time series data up to 30 days long. Through comparative experiments, it was found that after adding the residual connection, the accuracy of the model in long-term prediction tasks (more than 15 days) increased by 31%, and the root mean square error decreased by 0.28.

[0170] To enhance the ability to identify abnormal events, an anomaly detection module based on VAE was designed in this application. This module consists of an encoder and a decoder, both of which use a three-layer fully connected network (with the dimension change of 256→128→64→32), and learn the normal distribution pattern of the data by minimizing the reconstruction error. In practical applications, when the reconstruction error at a certain time point is detected to exceed 3 standard deviations of the normal distribution, the system will trigger an anomaly warning. This mechanism successfully predicted two severe pollution events in the summer of 2023, with the early warning time reaching 72 hours, 2.5 days earlier than the traditional method.

[0171] The overall performance of the model was verified through multiple metrics: during the regular monitoring period, the average accuracy of time series prediction reached 91%, exceeding the baseline model by 15 percentage points; in terms of anomaly event detection, the recall rate and precision rate reached 88% and 85% respectively, and the false negative rate was controlled below 7%. Especially when dealing with the pollution accumulation process under complex weather conditions, the model demonstrated excellent prediction ability and accurately characterized the evolution trend of pollutant concentration.

[0172] Through comparative experiments in multiple industrial parks in the Yangtze River Delta region, the time series feature extraction scheme of this application demonstrated significant advantages. Compared with the traditional single LSTM model, the prediction accuracy increased by 36%, the computational efficiency increased by 28%, and the memory occupancy decreased by 45% at the same time. These experimental data fully prove the innovative value and practicality of this scheme in the field of environmental risk warning.

[0173] The prediction model adopts an ensemble learning strategy, combining the outputs of multiple predictors. In addition to the main BiLSTM network, it also includes a temporal convolutional network (TCN) and a Prophet model as supplements. TCN captures multi-scale time patterns through causal convolution and dilated convolution, and the Prophet model is specifically designed to handle seasonal and periodic changes. The prediction results of these three models are fused through an adaptive weight mechanism, and the weights are dynamically adjusted according to the performance of each model on the validation set. The prediction uncertainty is estimated by the Monte Carlo dropout method, providing a confidence interval for each predicted value and improving the reliability of the prediction results.

[0174] S5.3.2: Construct a multi - layer early warning network based on the risk prediction sequence, identify high - risk areas and key influencing factors through the attention mechanism, and obtain the early warning evaluation results;

[0175] The multi - layer early warning network adopts a hierarchical risk assessment structure. The first layer is the spatial attention layer, which uses the self - attention mechanism to calculate the mutual influence between different regions. The attention score is calculated by scaled dot - product: A = softmax(QK^T / √d), where Q and K are the query and key matrices, and d is the feature dimension. The second layer is the temporal attention layer, which focuses on critical moments in the time series, especially periods when risks are rising or accumulating rapidly. The third layer is the feature attention layer, which identifies environmental factors that have an important impact on risk evolution. The attention weights are calculated through a learnable parameter matrix and normalized by the softmax function.

[0176] The identification of high - risk areas adopts a method based on graph neural networks. A spatial relationship graph is constructed, where nodes represent different regions and edges represent the spatial associations between regions. Neighborhood information is aggregated through graph convolution operations to identify risk propagation paths and aggregation areas. The identification of key influencing factors combines SHAP value analysis and gradient integration methods to quantify the contribution of different environmental factors to the risk state. The early warning evaluation results include information in multiple dimensions such as risk levels, spatial distributions, temporal evolution trends, and key influencing factors.

[0177] In the field of environmental risk early warning, traditional methods are often limited to single - dimension analysis and are difficult to comprehensively grasp the multi - dimensional evolution characteristics of risks. To address this problem, this application proposes an innovative multi - layer early warning network architecture that enables collaborative analysis in three dimensions: space, time, and features.

[0178] In the practice of Suzhou Industrial Park, the spatial attention layer first constructs a 25×25 attention matrix corresponding to the spatial association relationship of 5 - kilometer × 5 - kilometer grids within the park. The query matrix Q and key matrix K are calculated through a learnable parameter matrix W (with a dimension of 256×256). Experiments show that this attention mechanism can accurately capture the spatial transmission law of pollutants. For example, during the heavy pollution period in winter 2023, it successfully identified the significant impact of the northwest - coming flow on the air quality of the park, and the correlation coefficient reached 0.82.

[0179] The temporal attention layer adopts an innovative multi - scale sliding window mechanism. For short - term risks (within 24 hours), fine - grained analysis at 1 - hour intervals is used; for medium - term risks (1 - 7 days), 6 - hour intervals are adopted; for long - term risks (7 - 30 days), daily averages are used for analysis. This multi - scale strategy enables the model to simultaneously focus on the risk evolution characteristics at different time scales. Experimental data shows that this mechanism improves the early identification rate of risk events to 85% and discovers potential risk hazards 36 hours earlier on average.

[0180] In the feature attention layer, the model calculates the importance weights of environmental factors through a three-layer fully connected network (with the dimension changing as 256→128→64). Combining with SHAP value analysis, the contributions of different factors to the risk status are quantified. Taking the ozone pollution event in the summer of 2023 as an example, the model identifies that air temperature (weight 0.35), NOx concentration (weight 0.28), and light intensity (weight 0.25) are the main influencing factors, and this result is highly consistent with the actual observations.

[0181] S5.3.3: Classify the early warning assessment results and perform spatial clustering to generate early warning information.

[0182] In S5.3.3, the early warning level classification adopts an adaptive threshold method. First, the distribution characteristics of risk values are analyzed through kernel density estimation (KDE), and then the Jenks natural breaks method is used to determine the level boundaries. To improve the stability of classification, a time smoothing mechanism is introduced to avoid frequent level fluctuations. The early warning levels are usually divided into four levels: level one (red, severe risk), level two (orange, high risk), level three (yellow, medium risk), and level four (blue, low risk). Each level corresponds to different risk characteristics and prevention and control measures.

[0183] The spatial clustering adopts an improved DBSCAN algorithm, which can adaptively identify risk aggregation areas with irregular shapes. The two key parameters of the algorithm: the neighborhood radius ε and the minimum number of points MinPts are determined through grid search optimization. To handle the ambiguity of boundary regions, a soft clustering mechanism is introduced, allowing boundary points to belong to multiple clusters simultaneously and assigning corresponding membership degrees. The clustering results are overlaid and analyzed with administrative divisions and natural geographical units to generate an early warning map that is easy to understand and operate.

[0184] The final early warning information includes: risk level distribution maps, high-risk area identification results, risk evolution trend maps, key influencing factor analysis reports, etc. The visualization of the early warning information adopts a multi-dimensional display method, including professional GIS layers, intuitive early warning symbols, and detailed analysis reports. Experimental verification shows that the prediction accuracy of this early warning information reaches more than 85%, and the average early warning lead time can reach 3 - 5 days, providing strong decision-making support for environmental risk prevention and control.

[0185] As Figure 7 shown, the embodiment of the present application also provides a device for monitoring the environmental impact of territorial space planning, including:

[0186] A data acquisition and preprocessing module 10, configured to acquire and preprocess at least two types of multi-source environmental monitoring data including satellite remote sensing images, meteorological station data, land cover data, and ecological environment monitoring data;

[0187] The spatio-temporal feature extraction module 20 is used to extract spatio-temporal features from the preprocessed multi-source environmental monitoring data by using a self-supervised spatio-temporal graph mask transfer attention network to obtain initial spatio-temporal features;

[0188] The feature fusion module 30 is used to perform fusion processing on the initial spatio-temporal features based on a multi-modal diffusion transformer to obtain fusion features;

[0189] The evaluation index calculation module 40 is used to calculate environmental impact evaluation indexes including land use change index, ecological damage degree, and environmental quality comprehensive index based on the fusion features;

[0190] The visualization and warning module 50 is used to generate dynamic monitoring results including spatio-temporal evolution maps, trend analysis charts, and warning information based on the preset visualization templates and warning rules according to the environmental impact evaluation indexes.

[0191] It should be noted that the specific details of each step of the above method for monitoring the environmental impact of territorial space planning are applicable to the corresponding modules, and will not be elaborated here.

[0192] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, advantages, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present application. In addition, the above disclosed specific details are only for the purposes of illustration and easy understanding, and not for limitation. The above details do not limit the present application to necessarily adopt the above specific details to implement.

[0193] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the form disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A method for monitoring the environmental impact of national land space planning, characterized in that: include: Collecting and preprocessing multi-source environmental monitoring data, wherein the multi-source environmental monitoring data includes at least two of satellite remote sensing images, meteorological station data, surface coverage data, and ecological environment monitoring data; The self-supervised spatiotemporal graph mask transfer attention network is used to extract spatiotemporal features from the preprocessed multi-source environmental monitoring data to obtain the initial spatiotemporal features. Based on a multimodal diffusion transformer, the initial spatiotemporal features are fused to obtain fused features; Calculating environmental impact assessment indicators based on the fusion features; Based on the preset visualization template and early warning rules, dynamic monitoring results including spatiotemporal evolution diagrams, trend analysis charts and early warning information are generated according to the environmental impact assessment indicators; The method of extracting spatiotemporal features from the preprocessed multi-source environmental monitoring data using a self-supervised spatiotemporal graph mask transfer attention network to obtain the initial spatiotemporal features includes: Based on the preprocessed multi-source environmental monitoring data, the monitoring area is divided into space-time unit nodes, spatial edges are established based on geographical location proximity, and time edges are established based on time series continuity to construct a space-time relationship graph; For the nodes of the spatiotemporal relationship graph, 15%-30% of the node features are randomly masked, and the masked node features are reconstructed through the surrounding node information for self-supervised pre-training to obtain a pre-trained attention network model; Input the spatiotemporal relationship graph into the pre-trained attention network model, calculate the attention weights between spatially adjacent nodes through the spatial attention layer, and model the long-range dependency between different time points through the temporal attention layer to obtain the node attention weights; Based on the node attention weight, the node features are iteratively updated, the node features in the spatial neighborhood are integrated, and the gated recurrent unit is used to process the time dimension feature evolution to obtain the initial spatiotemporal features; The fusing process of the initial spatiotemporal features based on the multimodal diffusion transformer to obtain the fused features includes: Performing feature encoding on the initial spatiotemporal features and adding modality type tags to obtain coded features with modality tags; Constructing a multi-layer diffusion transformer including a diffusion attention layer and a cross-modal interaction block, and inputting the modality-labeled encoded features into the multi-layer diffusion transformer to obtain initial transformation features; Gradually adding Gaussian noise to the initial transformation features through forward diffusion, and then reconstructing the features through reverse diffusion to obtain reconstructed features; The reconstruction features are aligned in time and space to obtain registration features, the modal reliability weights of the registration features are calculated, and the features are dynamically fused based on the modal reliability weights to obtain the fused features.

2. The method for monitoring the environmental impact of land space planning according to claim 1, characterized in that: The calculating of the environmental impact assessment index based on the fusion feature includes: Based on the fusion features, an attention feature extraction network is constructed, and the environmental factor change features are extracted through the spatial attention mechanism and the temporal attention mechanism to obtain the environmental factor change sequence; Performing spatiotemporal autocorrelation analysis and change trend analysis on the environmental factor change sequence, calculating the environmental pressure index and ecological response index through a multi-layer perceptron, and obtaining the initial environmental impact assessment results; The initial environmental impact assessment results are segmented by adaptive thresholds and fused with multi-scale features to calculate environmental impact assessment indicators, wherein the environmental impact assessment indicators include environmental stress, ecological vulnerability and environmental quality index.

3. The method for monitoring the environmental impact of land space planning according to claim 2, characterized in that: The temporal and spatial autocorrelation analysis and change trend analysis of the environmental factor change sequence are performed, and the environmental pressure index and ecological response index are calculated by a multi-layer perceptron to obtain the initial environmental impact assessment results, including: Inputting the environmental factor change sequence into the spatiotemporal autocorrelation network, calculating the spatial aggregation characteristics and temporal evolution characteristics of the environmental factors through the local spatial autocorrelation operator and the global temporal autocorrelation operator, and obtaining the spatiotemporal correlation characteristics; The spatiotemporal correlation features are nonlinearly mapped and feature extracted by a multi-layer perceptron, the spatiotemporal distribution of the environmental pressure index and the ecological response index are calculated, and a pressure response distribution map is obtained; Based on the pressure response distribution diagram, an initial environmental impact assessment result is calculated by an adaptive weight fusion algorithm.

4. The method for monitoring the environmental impact of land space planning according to claim 3 is characterized in that: The calculating of the initial environmental impact assessment result based on the pressure response distribution map by an adaptive weight fusion algorithm includes: Performing multi-scale feature decomposition on the pressure response distribution map, extracting environmental pressure features and ecological response features at different scales through wavelet transform, and obtaining a multi-scale feature set; Building an adaptive weight network based on the multi-scale feature set, calculating weight coefficients of features of different scales through an attention mechanism, and obtaining a feature weight matrix; The multi-scale feature set and the feature weight matrix are weightedly fused, and an initial environmental impact assessment result is calculated through a deep neural network.

5. The method for monitoring the environmental impact of land space planning according to claim 2, characterized in that: The initial environmental impact assessment result is segmented by adaptive threshold and fused with multi-scale features to calculate environmental impact assessment indicators, including: Based on the initial environmental impact assessment results, an adaptive threshold network is constructed, and the grading thresholds of environmental stress, ecological vulnerability and environmental quality index are determined through multi-layer iterative optimization to obtain the assessment index thresholds; The initial environmental impact assessment results are graded according to the assessment index threshold, and multi-scale spatial features are extracted through a deep convolutional network to obtain a graded feature map; The hierarchical feature map is subjected to feature fusion and spatial clustering to calculate the environmental impact assessment index.

6. The method for monitoring the environmental impact of land space planning according to claim 5, characterized in that: The dynamic monitoring results including spatiotemporal evolution diagrams, trend analysis charts and warning information are generated based on the preset visualization templates and warning rules according to the environmental impact assessment indicators, including: Inputting the environmental impact assessment index into a dynamic visualization network, generating an environmental impact spatiotemporal evolution diagram through a spatiotemporal interpolation algorithm and gradient color mapping, and obtaining a visualization base layer; Performing multi-dimensional feature analysis on the visualization basic layer, generating an environmental change trend chart through a trend prediction model, and obtaining a trend analysis result; Based on the trend analysis results, an early warning assessment model is constructed, and the environmental risk level and spatial distribution are identified through a deep learning algorithm to generate early warning information.

7. The method for monitoring the environmental impact of land space planning according to claim 6, characterized in that: The early warning assessment model is constructed based on the trend analysis results, and the environmental risk level and spatial distribution are identified through a deep learning algorithm to generate early warning information, including: Extracting time series features from the trend analysis results, predicting the environmental risk evolution trend through a recurrent neural network, and obtaining a risk prediction sequence; Based on the risk prediction sequence, a multi-layer early warning network is constructed, and high-risk areas and key influencing factors are identified through an attention mechanism to obtain early warning evaluation results; The warning assessment results are graded and spatially clustered to generate warning information.

8. A national land space planning environmental impact monitoring device, characterized in that: include: A data collection and preprocessing module, used for collecting and preprocessing at least two types of multi-source environmental monitoring data including satellite remote sensing images, meteorological station data, surface coverage data and ecological environment monitoring data; The spatiotemporal feature extraction module is used to extract spatiotemporal features from the preprocessed multi-source environmental monitoring data using a self-supervised spatiotemporal graph mask transfer attention network to obtain initial spatiotemporal features; A feature fusion module, used for fusing the initial spatiotemporal features based on a multimodal diffusion transformer to obtain fused features; An evaluation index calculation module, used to calculate environmental impact assessment indicators including land use change index, ecological damage degree and environmental quality comprehensive index based on the fusion features; A visual early warning module, for generating dynamic monitoring results including spatiotemporal evolution diagrams, trend analysis charts and early warning information according to the environmental impact assessment indicators based on preset visual templates and early warning rules; The method of extracting spatiotemporal features from the preprocessed multi-source environmental monitoring data using a self-supervised spatiotemporal graph mask transfer attention network to obtain the initial spatiotemporal features includes: Based on the preprocessed multi-source environmental monitoring data, the monitoring area is divided into space-time unit nodes, spatial edges are established based on geographical location proximity, and time edges are established based on time series continuity to construct a space-time relationship graph; For the nodes of the spatiotemporal relationship graph, 15%-30% of the node features are randomly masked, and the masked node features are reconstructed through the surrounding node information for self-supervised pre-training to obtain a pre-trained attention network model; Input the spatiotemporal relationship graph into the pre-trained attention network model, calculate the attention weights between spatially adjacent nodes through the spatial attention layer, and model the long-range dependency between different time points through the temporal attention layer to obtain the node attention weights; Based on the node attention weight, the node features are iteratively updated, the node features in the spatial neighborhood are integrated, and the gated recurrent unit is used to process the time dimension feature evolution to obtain the initial spatiotemporal features; The fusing process of the initial spatiotemporal features based on the multimodal diffusion transformer to obtain the fused features includes: Performing feature encoding on the initial spatiotemporal features and adding modality type tags to obtain coded features with modality tags; Constructing a multi-layer diffusion transformer including a diffusion attention layer and a cross-modal interaction block, and inputting the modality-labeled encoded features into the multi-layer diffusion transformer to obtain initial transformation features; Gradually adding Gaussian noise to the initial transformation features through forward diffusion, and then reconstructing the features through reverse diffusion to obtain reconstructed features; The reconstruction features are aligned in time and space to obtain registration features, the modal reliability weights of the registration features are calculated, and the features are dynamically fused based on the modal reliability weights to obtain the fused features.

Citation Information

Patent Citations

  • Ecological environment monitoring method and system based on ecological function data analysis

    CN118779644A

  • Equipment remaining service life prediction based on self-supervised graph attention time sequence graph reasoning

    CN118885980A