A multi-source heterogeneous environment monitoring data processing method and system

By employing a multi-source heterogeneous environmental data processing method, the problem of data fragmentation was solved, enabling efficient and accurate prediction of environmental pollution spread. This improved the sensitivity and accuracy of the model and provided physical interpretability to support environmental source tracing and control decisions.

CN121682141BActive Publication Date: 2026-04-17陕西省环境监测中心站
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
陕西省环境监测中心站
Filing Date
2026-02-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies suffer from fragmented data processing in environmental pollution diffusion prediction and assessment, resulting in low information utilization, difficulty in meeting real-time emergency response needs, lack of physical interpretability, and inability to effectively combine industrial data analysis with the physical dynamics of pollutant diffusion.

Method used

By collecting multi-source heterogeneous environmental data, performing protocol parsing and spatiotemporal grid alignment, a unified spatiotemporal grid dataset is constructed. Multi-scale convolution operations are used to extract spatiotemporal correlation features, generating multi-layer spatiotemporal feature maps. A comprehensive environmental state vector is generated through weighted superposition and fusion encoding. Input parameter regression mapping model is used to analyze diffusion dynamics parameters, and a Lagrange particle tracking model is combined to iteratively calculate future pollution diffusion paths.

Benefits of technology

It enables the keen detection of sudden environmental pollution events, improves the agility and accuracy of the model, provides physical interpretability and high-precision prediction of future pollution trends, and supports environmental source tracing and control decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121682141B_ABST
    Figure CN121682141B_ABST
Patent Text Reader

Abstract

This invention relates to the field of environmental monitoring data processing technology, and discloses a method and system for processing multi-source heterogeneous environmental monitoring data. The method includes collecting multi-source heterogeneous environmental data and calculating its fluctuation characteristics to obtain environmental dynamic change indicators and a unified spatiotemporal grid dataset; extracting features from the unified spatiotemporal grid dataset to obtain a multi-layer spatiotemporal feature map; weightedly fusing and encoding the environmental dynamic change indicators and the multi-layer spatiotemporal feature map to obtain a comprehensive environmental state vector; analyzing the physical characteristics of the comprehensive environmental state vector to obtain a diffusion dynamics parameter set; calibrating and updating the diffusion dynamics parameter set to obtain a dynamic pollution diffusion path; and performing spatial gradient extrapolation along the dynamic pollution diffusion path to obtain a final pollution trend prediction result. This method can deeply integrate deep features with diffusion patterns, achieving high-precision dynamic prediction of the future spatiotemporal distribution of pollutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring data processing technology, and in particular to a method and system for processing multi-source heterogeneous environmental monitoring data. Background Technology

[0002] Currently, with the deep integration of industrial IoT technology and smart environmental protection construction, how to effectively utilize industrial data analysis and industrial data mining technology to extract core features that can accurately characterize the diffusion patterns of pollutants from these complex data containing spatiotemporal dimensions and different physical attributes has become the key to achieving refined environmental supervision and pollution source tracing.

[0003] In existing technologies, the prediction and assessment of environmental pollution diffusion mainly rely on traditional numerical air quality models (such as CALPUFF and AERMOD) or interpolation algorithms based on simple statistics. These technologies typically collect meteorological parameters directly as input, perform offline iterative calculations using fluid dynamics equations, or use methods such as inverse distance weighting to statically smooth discrete monitoring point data, thereby generating a regional pollutant concentration distribution map and formulating corresponding control measures. Purely physical mechanism models have extremely high computational complexity, making them difficult to meet the needs of real-time emergency response, while purely data-driven statistical models lack physical interpretability and are prone to black-box prediction bias. Faced with complex and ever-changing real-world environments, existing technologies struggle to deeply couple the statistical laws derived from industrial data analysis with the physical dynamics of pollutant diffusion, resulting in insufficient accuracy in predicting future pollutant diffusion paths and concentration gradients, severely restricting the scientific rigor and timeliness of environmental regulatory decisions.

[0004] Existing technologies suffer from fragmented data processing, leading to low information utilization. Summary of the Invention

[0005] This invention provides a method and system for processing multi-source heterogeneous environmental monitoring data to solve the problem of low information utilization caused by the fragmentation of data processing in the prior art.

[0006] In a first aspect, to address the aforementioned technical problems, this invention provides a method for processing multi-source heterogeneous environmental monitoring data, comprising: collecting multi-source heterogeneous environmental data; performing protocol parsing on the multi-source heterogeneous environmental data and aligning it with spatiotemporal grids to construct a unified spatiotemporal grid dataset; and calculating environmental dynamic change indicators characterizing data fluctuations based on the unified spatiotemporal grid dataset; performing multi-scale convolution operations on the unified spatiotemporal grid dataset, extracting the spatiotemporal correlation features of each grid unit under different receptive fields using convolution kernels of different sizes, and generating a multi-layer spatiotemporal feature map; using weighted superposition, fusing and encoding the environmental dynamic change indicators with the multi-layer spatiotemporal feature map to generate a comprehensive environmental state vector containing real-time dynamic environmental information; and inputting the comprehensive environmental state vector into a preset parameter regression mapping model to analyze and obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion, and diffusion dynamics... The diffusion kinetics parameter set includes the dominant diffusion direction, transmission velocity, and source strength attenuation coefficient. Based on the diffusion kinetics parameter set, a current linear motion equation is constructed, and the theoretical predicted coordinates of the pollution center at the current moment are calculated. The Euclidean distance between the theoretical predicted coordinates and the coordinates of the high-concentration center actually observed in the unified spatiotemporal grid dataset is calculated as the state deviation. Using a preset feedback gain coefficient, the transmission velocity and dominant diffusion direction are corrected in reverse according to the state deviation to obtain an updated diffusion kinetics parameter set. The updated parameter set is substituted into a preset Lagrange particle tracking model to iteratively calculate the center position coordinates at future moments until a preset prediction duration threshold is reached, generating a dynamic pollution diffusion path. A preset concentration attenuation function is applied along the dynamic pollution diffusion path to perform spatial gradient extrapolation, calculate the spatiotemporal concentration distribution at future moments, and obtain the final pollution trend prediction result.

[0007] Secondly, this invention provides a system for processing multi-source heterogeneous environmental monitoring data, comprising: a data preprocessing module for collecting multi-source heterogeneous environmental data, performing protocol parsing on the multi-source heterogeneous environmental data, aligning the multi-source heterogeneous environmental data with spatiotemporal grids, constructing a unified spatiotemporal grid dataset, and calculating environmental dynamic change indicators characterizing data fluctuations based on the unified spatiotemporal grid dataset; a multi-scale feature extraction module for performing multi-scale convolution operations on the unified spatiotemporal grid dataset, extracting the spatiotemporal correlation features of each grid unit under different receptive fields using convolution kernels of different sizes, and generating a multi-layer spatiotemporal feature map; a state fusion encoding module for using weighted superposition to fuse and encode the environmental dynamic change indicators with the multi-layer spatiotemporal feature map, generating a comprehensive environmental state vector containing real-time dynamic environmental information; and a physical parameter inversion module. The system is divided into four modules: a dynamic path correction module and a trend prediction module. The dynamic path correction module is used to input the comprehensive environmental state vector into a preset parametric regression mapping model and analyze it to obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion. The set of diffusion dynamic parameters includes the dominant diffusion direction, transmission velocity, and source strength attenuation coefficient. The dynamic path correction module is used to use the diffusion dynamic parameter set to deduce the theoretical prediction state at the current moment, calculate the state deviation between the theoretical prediction state and the current moment's observation value in the unified spatiotemporal grid dataset, calibrate and update the diffusion dynamic parameter set based on the state deviation, and use the updated parameter set to drive a preset diffusion evolution model to generate a dynamic pollution diffusion path for the future. The trend prediction module is used to apply a preset concentration attenuation function to perform spatial gradient deduction along the dynamic pollution diffusion path, calculate the spatiotemporal concentration distribution at future moments, and obtain the final pollution trend prediction result.

[0008] Compared with the prior art, the present invention has the following beneficial effects:

[0009] (1) This invention significantly improves the model's ability to sensitively capture sudden environmental pollution events by introducing a weighted fusion mechanism of environmental dynamic change indicators and multi-scale features. Existing technologies often only focus on static concentration distribution, easily ignoring instantaneous and drastic fluctuations in the environmental field. This invention extracts environmental dynamic change indicators by calculating the variance of grid cells and embeds them into the spatiotemporal feature map extracted by a multi-scale convolutional network using an attention mechanism. This fusion strategy enables the model to automatically focus on "active areas" with drastic data fluctuations and high potential risks during the feature encoding stage, thereby effectively solving the problems of delayed response and insufficient feature extraction of traditional models to sudden point source emissions or rapid diffusion caused by gusts, and realizing a refined perception of dynamic environmental changes.

[0010] (2) This invention achieves bidirectional coupling between "data and physics" by constructing a parametric regression mapping model, breaking through the bottleneck of the lack of interpretability in pure data-driven models. Existing deep learning models are usually "black boxes," directly outputting prediction results without explaining the intermediate processes. This invention innovatively maps the latent space features of deep neural networks into a set of diffusion dynamic parameters with clear physical meaning (diffusion dominance direction, transmission speed, and source strength attenuation coefficient). This design not only retains the powerful nonlinear fitting ability of deep learning but also endows the model with physical interpretability, enabling regulators to intuitively understand the intrinsic mechanism of pollution diffusion by observing parameter changes (such as sudden changes in wind speed parameters), providing a scientific and transparent physical basis for environmental source tracing and control decisions.

[0011] (3) This invention eliminates the accumulated error in long-term time-series extrapolation through a closed-loop calibration mechanism based on trajectory deviation dynamic parameters, significantly improving the accuracy and robustness of future predictions. Traditional diffusion models operate in an open loop once initial conditions are set, and the error will continuously amplify over time. This invention introduces a feedback adjustment mechanism similar to that in control theory, which calculates the deviation between the theoretically predicted state and the current actual observation value in real time, and uses this deviation to correct key dynamic parameters such as transmission speed and direction. This closed-loop logic of "prediction-correction-re-prediction" enables the system to adaptively offset the effects of sensor noise and unmodeled environmental disturbances, ensuring that the generated dynamic pollution diffusion path always closely follows the real physical evolution trajectory, thereby achieving high-precision prediction of future pollution trends. Attached Figure Description

[0012] Figure 1 This is a schematic flowchart of a method for processing multi-source heterogeneous environmental monitoring data provided in an embodiment of the present invention;

[0013] Figure 2 This is a schematic diagram of a multi-source heterogeneous environmental monitoring data processing system provided in an embodiment of the present invention. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Reference Figure 1 This invention provides a method for processing multi-source heterogeneous environmental monitoring data, comprising the following steps:

[0016] S11, Collect multi-source heterogeneous environmental data, perform protocol parsing on the multi-source heterogeneous environmental data, perform spatiotemporal grid alignment, construct a unified spatiotemporal grid dataset, and calculate environmental dynamic change indicators characterizing data fluctuation characteristics based on the unified spatiotemporal grid dataset.

[0017] S12, perform multi-scale convolution operation on the unified spatiotemporal grid dataset, and use convolution kernels of different sizes to extract the spatiotemporal correlation features of each grid unit under different receptive fields to generate multi-layer spatiotemporal feature maps.

[0018] S13, using weighted superposition, the environmental dynamic change index and the multi-layer spatiotemporal feature map are fused and encoded to generate a comprehensive environmental state vector containing real-time dynamic environmental information.

[0019] S14, Input the comprehensive environmental state vector into a preset parameter regression mapping model, and analyze to obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion;

[0020] S15, using the diffusion dynamics parameter set to deduce the theoretical prediction state at the current moment, calculating the state deviation between the theoretical prediction state and the current moment observation value in the unified spatiotemporal grid dataset, calibrating and updating the diffusion dynamics parameter set based on the state deviation, and using the updated parameter set to drive the preset diffusion evolution model to generate a future-oriented dynamic pollution diffusion path.

[0021] S16, along the dynamic pollution diffusion path, a preset concentration decay function is applied to perform spatial gradient extrapolation, calculate the spatiotemporal concentration distribution at future times, and obtain the final pollution trend prediction result.

[0022] In step S11, multi-source heterogeneous environmental data is collected, protocol parsing is performed on the multi-source heterogeneous environmental data, and spatiotemporal grid alignment is carried out to construct a unified spatiotemporal grid dataset. Based on the unified spatiotemporal grid dataset, environmental dynamic change indicators characterizing data fluctuation characteristics are calculated, including:

[0023] Based on the preset protocol parsing rules, fields are extracted from the raw data packets collected from devices from different manufacturers to obtain standard format data containing latitude and longitude, timestamps and monitoring values;

[0024] The standard format data is mapped to a spatiotemporal grid coordinate system with a preset resolution using inverse distance weighted interpolation. The grid cells of missing data are then interpolated to complete the data, resulting in a unified spatiotemporal grid dataset.

[0025] A time sliding window of a preset length is constructed, and the variance value of the monitoring data in each grid cell is calculated based on the time sliding window. The variance value is then used as the environmental dynamic change index that characterizes the data fluctuation characteristics.

[0026] It should be noted that extracting fields from raw data packets collected from devices from different manufacturers according to pre-defined protocol parsing rules is a key step in solving the device heterogeneity problem. This operation first establishes a Device Profile Library at the edge gateway layer. This library explicitly defines the supported communication protocol types (such as Modbus-RTU, MQTT, CoAP) and the byte mapping rules for each device's packets. When the system receives a raw binary or hexadecimal packet stream, it first indexes the corresponding profile file based on the device ID in the packet header. Subsequently, it uses bitwise operations to extract data segments according to predefined offsets and lengths. Specifically, the predefined offsets and lengths are determined based on the hardware communication protocol datasheet provided by the device manufacturer. The system identifies the starting byte position of the target physical quantity (such as PM2.5 value) in the data frame payload as the offset by parsing the register address mapping table in the datasheet, identifies the number of bytes it occupies as the length, and pre-configures these values ​​in the device profile file. Finally, the truncated values ​​are restored to physically meaningful floating-point numbers by applying preset scaling factors and offsets, and the parsed longitude, latitude, UTC timestamps, and environmental monitoring values ​​(such as PM2.5 concentration and temperature) are encapsulated into a unified JSON key-value pair format, i.e., the standard format data.

[0027] It is worth noting that the preset scaling factor and offset are determined based on linear regression analysis of the sensor's factory calibration data or laboratory calibration data. Specifically, the construction process involves building a set of standard physical quantity input sequences covering the sensor's measurement range under standard experimental conditions. (e.g., standard gas concentration), and record the corresponding raw digital output sequence of the sensor. The linear transformation equation was fitted using the least squares method. The slope calculated therefrom That is, the scaling factor, the intercept. This is the offset. The system only saves the set of parameters to the device description file library when the goodness-of-fit (R-squared) is greater than a preset linear fit confidence threshold (e.g., 0.999), to ensure the accuracy of physical quantity reconstruction. The linear fit confidence threshold is determined based on the requirements of the International Organization of Legal Metrology (OIML) R76 standard for the accuracy class of precision instruments, and is obtained by calculating the minimum correlation coefficient required to meet a specific confidence interval (e.g., 99%).

[0028] It should be noted that the inverse distance weighted interpolation (IDW) method is used to map the standard format data to a spatiotemporal grid coordinate system of a preset resolution, aiming to construct a spatially continuous data field. This operation first divides the monitoring area into... For a regular grid, for each blank grid cell without direct observation data, all valid observation points within its search radius are searched. Then, the weighted average of the values ​​of each valid observation point is calculated as the interpolation result for that grid cell. The weight of each observation point is inversely proportional to the power of its Euclidean distance to the grid center; the closer the distance, the greater the weight.

[0029] It is worth noting that the preset resolution (e.g., 10 meters) The determination of the 10-meter range is based on spatial autocorrelation analysis of environmental parameters in the monitoring area. The semivariogram of historical monitoring data is calculated, and a variogram curve is fitted. The range at which this curve reaches the sill value is extracted; this range reflects the spatial correlation distance of the environmental data. According to the sampling theorem, the grid resolution is set to 1 / 2 to 1 / 3 of this range to ensure that the grid can capture subtle spatial variations. The power parameter (usually set to 2) is determined based on cross-validation optimization. Different parameters are tested in the historical dataset. The root mean square prediction error (RMSE) at a given value is selected to minimize the error. The value is used as a preset parameter.

[0030] It should be noted that a time sliding window of a preset length is constructed, and the variance of the monitoring data within each grid cell is calculated based on this time sliding window. This aims to quantify the degree of fluctuation in environmental data over a short period of time. For any grid cell, the system maintains a time sliding window of a length of... The data is processed using a first-in, first-out (FIFO) data queue. Whenever new data is written or interpolated, the queue is updated. First, the arithmetic mean of all monitored data in the current queue is calculated. Then, the difference between each monitored data in the queue and the arithmetic mean is calculated and squared. Finally, the squares of all differences are summed and divided by the total length of the data in the queue. The result is used as the environmental dynamic change index characterizing the data fluctuation characteristics.

[0031] It is worth noting that the length of the time sliding window The determination of the window length (e.g., 30 minutes) is based on temporal spectral analysis of historical environmental data. By performing a Fast Fourier Transform (FFT) on the historical data, the main frequency range of environmental noise (such as dust fluctuations caused by gusts of wind) is identified. Choosing a window length that covers at least one complete noise cycle ensures that the calculated variance effectively reflects the dynamic disturbances of the local microenvironment, rather than long-term trend changes.

[0032] More specifically, the time sliding window length can be selected from long-term historical monitoring data of the monitoring area under typical meteorological conditions. This data undergoes detrending and standardization preprocessing. Then, a Fast Fourier Transform (FFT) is applied to the series to obtain a power spectral density (PSD) map. Analyzing the PSD map identifies peak frequencies with significantly higher power than the background level (typically corresponding to short-term fluctuations caused by gusts, traffic sources, etc.). Finally, the time window length is set to a preset period multiple divided by the peak frequency. The preset period multiple is typically greater than or equal to 2 to ensure the window can capture the complete periodic characteristics of the fluctuation component. For example, if the identified main noise frequency is 1 hour... -1 (With a period of 1 hour), if the period is multiplied by 2, then the time sliding window length is set to 2 hours. This method ensures that the variance calculation can effectively separate short-term dynamic fluctuations from long-term trends.

[0033] For example, the system receives a Modbus message "01 03 04 01F4..." from a PM2.5 detector. The protocol parser calls the parsing rules based on the ID "01", extracts the hexadecimal "01 F4", converts it to decimal 500, multiplies it by a scaling factor of 0.1, obtains the monitoring value of 50.0 µg / m³, and generates a JSON record by combining it with GPS data. Subsequently, during the gridding process, there are two observation points around a blank grid, at distances of 2 meters and 4 meters respectively, with observation values ​​of 50 and 60 respectively. If... The weights are 1 / 4 = 0.25 and 1 / 16 = 0.0625, respectively. The interpolation result is (0.25 × 50 + 0.0625 × 60) / (0.25 + 0.0625) + 16.25 / 0.3125 = 52. Finally, the five interpolated data points for this grid over the past 30 minutes are [50, 52, 51, 53, 52], and their variance is calculated to be 1.04. This value is output as the environmental dynamic change index for this grid at the current moment.

[0034] In step S12, multi-scale convolution operations are performed on the unified spatiotemporal grid dataset. Convolution kernels of different sizes are used to extract the spatiotemporal correlation features of each grid unit under different receptive fields, generating a multi-layer spatiotemporal feature map, including:

[0035] Construct a feature extraction network containing multiple parallel convolutional branches, where each branch is configured with dilated convolutional kernels with different dilation rates;

[0036] The unified spatiotemporal grid dataset is simultaneously input into the multiple parallel convolutional branches to extract local feature maps at different receptive field scales.

[0037] The local feature maps output from each branch are concatenated along the channel dimension, and a 1×1 convolutional layer is used for channel dimensionality reduction and feature integration to generate the multi-layer spatiotemporal feature map.

[0038] It should be noted that the feature extraction network containing multiple parallel convolutional branches is constructed using the Atrous Spatial Pyramid Pooling (ASPP) architecture. This network consists of 3 to 5 parallel convolutional layers, each branch using the same kernel size (e.g., 3×3) but configured with different dilation rates. Atrous convolution inserts dilation between the weights of the convolutional kernel... To expand the receptive field, use a zero value. The calculation logic for the equivalent receptive field size, denoted as the dilation rate, is as follows: first, calculate the difference between the convolution kernel size and one; multiply this difference by the dilation rate; and finally, add one to the product to obtain the equivalent receptive field size. This design allows the network to capture both the microscopic diffusion characteristics of adjacent grids and the macroscopic transmission characteristics across multiple grids without increasing the number of parameters or computational cost.

[0039] It is worth noting that the specific values ​​of the different expansion rates (e.g.) The determination of the diffusion radius is based on the statistical distribution analysis of historical pollutant diffusion radii. A large amount of historical pollution event data is collected to calculate the maximum and minimum diffusion distances of pollutants within a unit time step (e.g., 30 minutes). These distances are then mapped to a grid coordinate system to obtain the diffusion span in grid cells. An expansion rate sequence is set based on this range to ensure coverage of both the minimum and maximum expansion rates, thereby guaranteeing that the feature extraction network can fully cover all possible diffusion scales of the pollutants.

[0040] It should be noted that inputting the unified spatiotemporal grid dataset into the multiple parallel convolution branches simultaneously means that the shape of the output of S11 is... tensors (where For batch size, For time steps, The grid dimensions are: The data is input into the network (number of channels). Each branch performs convolution operations independently and outputs data with the same spatial resolution. However, these are local feature maps representing semantic information at different scales. For example, the branch with an expansion rate of 1 extracts the local abrupt change features of a single point and its 8-neighborhood (such as sudden point source emissions), while the branch with an expansion rate of 6 extracts the background concentration gradient features over a large area (such as regional haze).

[0041] It should be noted that the local feature maps output by each branch are concatenated along the channel dimension, and a 1×1 convolutional layer is used for channel dimensionality reduction and feature integration, aiming to fuse multi-scale information and control model complexity. First, a tensor concatenation operation is used to merge the feature maps output by each branch along the channel axis, resulting in a high-dimensional feature tensor; subsequently, the following steps are applied... The convolutional kernel performs a linear weighted combination on the high-dimensional tensor. This operation is mathematically equivalent to performing a fully connected weighted sum on features at different scales. By learning the weight parameters, it automatically determines which scales of features are more important to the current prediction task (e.g., increasing the weight of large-scale features when the wind speed is high), thereby generating the multi-layer spatiotemporal feature map containing rich contextual information.

[0042] It is worth noting that the determination of the number of output channels of the 1×1 convolutional layer (i.e., the dimension after feature integration) is based on a balance analysis between model compression ratio and information entropy loss. By testing the model prediction accuracy and inference latency under different numbers of output channels on the validation set, a precision-latency Pareto curve is plotted, and the number of channels corresponding to the inflection point of the curve is selected as the preset parameter to minimize computational overhead while ensuring feature expressiveness.

[0043] For example, one of the inputs A grid dataset (containing three channels: PM2.5, temperature, and wind speed). The network contains three parallel branches with expansion rates of 1, 2, and 5, respectively. Branch 1 ( Extract local features and output shape as Feature map; Branch 2 ( Extract mesoscale features and output shape as Feature map; Branch 3 ( Extract large-scale features and output shape as The feature map. After concatenation along the channel dimension, the resulting shape is... The tensor. Finally, through The convolutional layer reduces its dimensionality to 64 channels, and the output shape is... The multi-layer spatiotemporal feature map, in which each pixel of the feature map incorporates multi-scale environmental information from local to regional levels.

[0044] In step S13, a weighted superposition method is used to fuse and encode the environmental dynamic change index with the multi-layer spatiotemporal feature map to generate a comprehensive environmental state vector containing real-time dynamic environmental information, including:

[0045] The environmental dynamic change index is broadcast and expanded in the spatial dimension so that its size is consistent with the spatial resolution of the multi-layer spatiotemporal feature map, thus obtaining the fluctuation feature matrix.

[0046] A preset feature weight matrix is ​​introduced, and the fluctuation feature matrix is ​​weighted element by element.

[0047] The weighted fluctuation feature matrix and the weighted multi-layer spatiotemporal feature map are added element by element to obtain the fused feature tensor. The feature tensor is then subjected to global average pooling to generate the comprehensive environmental state vector.

[0048] It should be noted that broadcasting and expanding the environmental dynamic change index in the spatial dimension is a key preprocessing step to solve the problem of dimensional mismatch in multimodal data. Since the environmental dynamic change index (i.e., the variance map) output in step S11 is typically a two-dimensional single-channel real matrix composed of height and width dimensions, while the multi-layer spatiotemporal feature map output in step S12 is a three-dimensional multi-channel real tensor composed of height, width, and number of channels, this operation expands it into a three-dimensional tensor with geometrically identical height, width, and channel dimensions to the multi-layer spatiotemporal feature map by copying the environmental dynamic change index an equal number of times along the channel dimension, resulting in the fluctuation feature matrix. This ensures that subsequent weighting and fusion operations can be performed at each corresponding voxel level.

[0049] It should be noted that the element-wise weighting of the fluctuation feature matrix and the weighting of the multi-layer spatiotemporal feature map using a pre-set feature weight matrix are implemented using an attention-based feature recalibration mechanism. The system loads two pre-set weight matrices with the exact same geometric dimensions as the fluctuation feature matrix: a fluctuation weight matrix and a feature weight matrix. Using the Hadamard product (element-wise multiplication), the value at each position in the fluctuation feature matrix is ​​multiplied by the value at the same position in the fluctuation weight matrix to obtain the weighted fluctuation feature matrix. Simultaneously, the value at each position in the multi-layer spatiotemporal feature map is multiplied by the value at the same position in the feature weight matrix to obtain the weighted multi-layer spatiotemporal feature map. This operation explicitly amplifies the contribution of high-fluctuation regions (i.e., areas of drastic environmental change) to the final state vector by adjusting the values ​​of the weight matrices, while suppressing low-correlation background noise features.

[0050] It is worth noting that the predetermined weight matrix was constructed based on sensitivity analysis of historical pollution diffusion cases. A large number of historical samples were collected, and the values ​​at different positions and channels in the input data were fine-tuned sequentially using the perturbation method. The gradient of its impact on the final pollutant concentration prediction error was observed. The average gradient magnitude of each feature dimension was calculated and normalized to... The interval is used to obtain a weight distribution matrix reflecting the importance of each feature. This matrix is ​​fixed during the system initialization phase and used as a preset parameter for the inference phase.

[0051] It should be noted that Global Average Pooling (GAP) on the feature tensor aims to compress the feature tensor containing spatial information into a compact state vector. For the fused three-dimensional real feature tensor consisting of height, width, and number of channels, the GAP operation calculates the arithmetic mean of all pixel values ​​in each feature channel plane. Specifically, for each feature channel, the pixel values ​​in the spatial plane formed by the height and width of that channel are summed, and the sum is divided by the total number of pixels in that spatial plane. This operation eliminates the dependence of features on specific spatial locations, generating a one-dimensional real vector with a length equal to the number of channels in the three-dimensional real feature tensor, i.e., the integrated environment state vector, which can be used as input to the subsequent regression model.

[0052] For example, suppose the feature map size is The number of channels is 64. (In grid coordinates) At this location, the environmental dynamics index (variance) is 1.5. After broadcast expansion, the value at this location is 1.5 across all 64 channels. The preset fluctuation weight matrix assigns a weight of 0.8 to the first channel at this location. The weighted fluctuation characteristic value is... Meanwhile, the original value of the spatiotemporal feature map of the first channel at this location is 2.0, with a corresponding feature weight of 0.5. The weighted feature value is... The value of the first channel of the fused feature tensor at this position is... Finally, for the whole The average value of the first channel plane is calculated, and the result is assumed to be 1.85. This value is the first component of the integrated environmental state vector.

[0053] In step S14, the comprehensive environmental state vector is input into a preset parametric regression mapping model to analyze and obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion, including:

[0054] The comprehensive environmental state vector is input into the preset parameter regression mapping model and processed by the nonlinear activation function of the hidden layer to be mapped to the physical parameter space;

[0055] The normalized direction vector is calculated using the Tanh activation function of the output layer as the dominant diffusion direction. The scalar value of the transmission speed and the source strength attenuation coefficient are calculated using the Sigmoid activation function and a preset physical extreme value range.

[0056] It should be noted that the preset parameter regression mapping model is constructed using a Multi-Layer Perceptron (MLP) architecture. This model consists of an input layer, three hidden layers, and an output layer. The dimension of the input layer is consistent with the dimension of the comprehensive environment state vector generated in step S13 (e.g., 64 dimensions). The three hidden layers contain 128, 64, and 32 neurons respectively, and are fully connected, with ReLU (Rectified Linear Unit) uniformly configured as the nonlinear activation function to capture the high-order nonlinear mapping relationship between features and physical parameters. The output layer contains four neurons, corresponding to the X and Y components of the diffusion direction, the normalized value of the transmission velocity, and the normalized value of the source strength attenuation coefficient.

[0057] It is worth noting that the training process of the pre-set parameter regression mapping model is based on supervised learning of historical feature-parameter pairs. First, a large number of historical time periods' comprehensive environmental state vectors are collected as input samples; second, the corresponding real physical parameters (direction, velocity, decay rate) at that moment are calculated using Kalman filtering and used as labels; finally, mean squared error (MSE) is used as the loss function, and backpropagation training is performed using the Adam optimizer until the model's loss converges on the validation set.

[0058] It should be noted that the calculation of the normalized direction vector as the dominant diffusion direction using the Tanh activation function of the output layer is achieved through vector synthesis and normalization operations. The first two neurons in the model's output layer, responsible for representing the direction, output the original values. The hyperbolic tangent function (Tanh) is used to map these two original values ​​to a value range of -1 to +1, thus forming the two components of the original direction vector. Subsequently, the original direction vector is normalized using the L2 norm. Specifically, the magnitude of the original direction vector is calculated, and each component value is divided by this magnitude to obtain a unit direction vector with a magnitude of one, which is the dominant diffusion direction.

[0059] It should be noted that the calculation of the scalar value of the transmission speed and the source strength attenuation coefficient using the Sigmoid activation function and a preset physical extreme value range is achieved using a linear interval mapping method. In the model's output layer, the two neurons responsible for representing the transmission speed and the source strength attenuation coefficient output the original values. First, the Sigmoid activation function maps these two original values ​​to a dimensionless value range between zero and one. Then, for the transmission speed, the difference between the maximum and minimum values ​​within the preset speed physical extreme value range is calculated. This difference is multiplied by the mapped dimensionless value of the transmission speed, and the product is added to the minimum value of the speed physical extreme value range to obtain the scalar value of the transmission speed. Similarly, for the source strength attenuation coefficient, the difference between the maximum and minimum values ​​within the preset attenuation coefficient physical extreme value range is calculated. This difference is multiplied by the mapped dimensionless value of the source strength attenuation coefficient, and the product is added to the minimum value of the attenuation coefficient physical extreme value range to obtain the source strength attenuation coefficient.

[0060] The output layer design of the parametric regression mapping model aims to simultaneously regress multiple physical parameters through a multi-task learning framework. During training, the model's total loss function is a weighted sum of the prediction losses of each parameter. The weights corresponding to the cosine similarity loss or mean square error loss of the direction vector, as well as the mean square error loss of velocity and attenuation coefficient, are tuned according to the convergence difficulty and importance of each parameter on the validation set. This design enables the model to simultaneously optimize the prediction of all physical parameters during backpropagation and implicitly learn the physical relationships between them (such as wind speed affecting attenuation rate), thereby generating a more physically consistent set of parameters.

[0061] It is worth noting that the preset physical extreme value range (e.g., velocity range) m / s, attenuation coefficient range The determination of the limit is based on the statistical extreme value analysis of historical meteorological and environmental data of the monitoring area. The maximum wind speed and the fastest dissipation rate of pollutants recorded in the area over the past 5 years are statistically analyzed, and the 99.9th percentile of their distribution is selected as the upper limit value, while the lower limit value is set to 0 or the minimum physically permissible value (such as the background attenuation rate).

[0062] For example, we input a comprehensive environmental state vector. The first two raw values ​​of the model's output layer are 0.8 and -0.5. Tanh calculations yield a horizontal component value of 0.66 and a vertical component value of -0.46. After normalization, the diffusion-dominant direction vector is approximately... The third raw value of the model's output layer is 0.0. After Sigmoid calculation, it becomes 0.5. If the preset velocity range is... m / s, then the transmission speed is m / s. The fourth raw value of the model output layer is -1.38. The Sigmoid function calculates it to be 0.2. If the preset attenuation coefficient range is... Then the source strength attenuation coefficient is The final output parameter set is the direction. The speed is 10 m / s and the attenuation coefficient is 0.2.

[0063] In step S15, the theoretically predicted state at the current moment is derived using the diffusion dynamics parameter set; the state deviation between the theoretically predicted state and the current moment's observation in the unified spatiotemporal grid dataset is calculated; the diffusion dynamics parameter set is calibrated and updated based on the state deviation; and the updated parameter set is used to drive a preset diffusion evolution model to generate a future-oriented dynamic pollution diffusion path, including:

[0064] Based on the diffusion dynamics parameter set, the current linear motion equation is constructed, and the theoretical predicted coordinates of the pollution center at the current moment are calculated.

[0065] The Euclidean distance between the theoretically predicted coordinates and the actual observed high-concentration center coordinates in the unified spatiotemporal grid dataset is calculated as the state deviation.

[0066] Using a preset feedback gain coefficient, the transmission speed and diffusion dominance direction are corrected in reverse according to the state deviation to obtain an updated set of diffusion dynamics parameters;

[0067] The updated parameter set is substituted into the preset Lagrange particle tracking model, and the center position coordinates at future moments are calculated iteratively until the preset prediction duration threshold is reached, thereby generating a dynamic pollution diffusion path.

[0068] It is worth noting that the preset prediction duration threshold (e.g., 2 hours) is determined based on error divergence rate analysis of historical prediction data. Several typical pollution diffusion events occurring in historical years are selected, and this system is used to make advance predictions of different durations (e.g., from 1 hour to 24 hours). The mean absolute percentage error (MAPE) for each prediction duration is calculated, and a duration-error evolution curve is plotted. The critical point where the error derivative (i.e., the error growth rate) on the curve undergoes a significant abrupt change (e.g., the time point when MAPE first exceeds 20%) is found, and this critical time point is set as the prediction duration threshold to ensure that the output path is within the reliability range.

[0069] It should be noted that the calculation of the theoretical predicted coordinates of the pollution center at the current moment, based on the diffusion dynamics parameter set and the current linear motion equations, is a deduction based on kinematic principles. The system first obtains the pollution center coordinates of the previous time step; then, combining the transport velocity scalar and the diffusion dominant direction vector obtained in step S14, it constructs a linear advection calculation logic. First, the transport velocity scalar is multiplied by the lateral and longitudinal components of the diffusion dominant direction vector to obtain the velocity components of the pollutant on the horizontal and vertical axes; next, these two velocity components are multiplied by the data acquisition time interval to calculate the lateral and longitudinal displacement distances of the pollutant center within the current time step; finally, the lateral and longitudinal displacement distances are accumulated and added to the corresponding lateral and longitudinal coordinate values ​​of the pollution center coordinates of the previous time step to obtain the theoretical predicted coordinates at the current moment.

[0070] It should be noted that calculating the Euclidean distance between the theoretically predicted coordinates and the coordinates of the high-concentration centers actually observed in the unified spatiotemporal grid dataset first requires extracting the actual high-concentration centers from the unified spatiotemporal grid dataset using the Gray-level Centroid Method. The grid dataset at the current moment is traversed, and all grid points with monitoring values ​​greater than a preset background threshold are selected. The weighted average values ​​of all selected grid points on the horizontal and vertical axes are calculated, where the weights are the pollutant concentration values ​​of each grid point. Specifically, the calculation logic is as follows: the pollutant concentration of each grid point is multiplied by its horizontal coordinate value and summed. The sum is then divided by the total pollutant concentration of all selected grid points to obtain the horizontal coordinate of the actual observation center. Similarly, the pollutant concentration of each grid point is multiplied by its vertical coordinate value and summed. The sum is then divided by the total pollutant concentration of all selected grid points to obtain the vertical coordinate of the actual observation center. Subsequently, the differences between the theoretically predicted coordinates and the actual observation center coordinates in the horizontal and vertical directions are calculated respectively. The square root of the sum of the squares of these two differences is taken as the scalar value of the state deviation. At the same time, a vector pointing to the actual observation center is constructed using the difference between the actual observation center and the theoretically predicted coordinates in the horizontal direction as the horizontal component and the difference in the vertical direction as the vertical component. This vector serves as the deviation vector of the state deviation.

[0071] It is worth noting that the preset background threshold is determined based on a statistical distribution analysis of historical background noise data of the monitoring area during non-pollution periods. All grid monitoring data from periods of "excellent" or "good" air quality in the area over the past year were collected to construct a probability density function (PDF) for the background concentration. Based on the assumption of a normal distribution, the mean background concentration plus three standard deviations was selected as the preset background threshold. This setting aims to effectively filter out the inherent electronic white noise of the sensor and non-sudden environmental background fluctuations, ensuring that the grid points involved in the centroid calculation strictly correspond to the effective coverage area of ​​the high-concentration pollution plume, thereby improving the robustness of center positioning.

[0072] It should be noted that the reverse correction of the transmission speed and diffusion dominance direction based on the state deviation using preset feedback gain coefficients is implemented using proportional feedback control logic. The system introduces two preset feedback gain coefficients for speed correction and direction correction, respectively. First, for the transmission speed, the product of the magnitude of the state deviation (i.e., Euclidean distance error) and the feedback gain coefficient for speed correction is calculated, and this product is divided by the data acquisition time interval to obtain the speed compensation value. This speed compensation value is then added to the current transmission speed scalar. Simultaneously, for the diffusion dominance direction, an intermediate vector for correction is constructed. The construction logic is as follows: the current transmission speed scalar is multiplied by the current diffusion dominance direction vector to obtain the current speed vector; the deviation vector contained in the state deviation is multiplied by the feedback gain coefficient for direction correction and divided by the time interval to obtain the deviation speed component; the current speed vector and the deviation speed component are vector-wise added to obtain the intermediate vector; finally, the intermediate vector is normalized (i.e., divided by the magnitude of the intermediate vector) to obtain a new unit vector as the updated diffusion dominance direction. This step feeds back the residual between the model predictions and physical observations into the dynamic parameters, eliminating the accumulated error of a purely data-driven model.

[0073] As another specific and easily implemented method for determining the feedback gain coefficient, it can be optimized through grid search and cross-validation based on historical data. Specifically, a large amount of historical time-series data is collected, and theoretical predictions are made at each time step using an uncalibrated initial parameter set. Then, within a preset candidate value range (e.g., [0.1, 1.0]), all combinations are traversed with a fixed step size (e.g., 0.1). For each set of coefficients, the entire calibration update process is simulated, and the average Euclidean distance error between the future short-term predicted path (as of the next moment) generated after calibration and the actual observed path is calculated. Finally, the coefficient that minimizes this average error is selected. and By combining the feedback gain coefficients as preset by the system, this method does not rely on complex filter design and directly optimizes the calibration effect through data-driven means.

[0074] It should be noted that, in this embodiment, the updated parameter set is substituted into a preset diffusion evolution model to generate a dynamic pollution diffusion path. This is achieved using the Lagrangian Particle Tracking Model (LPTLM). This model treats pollutants as a set of virtual particle packets. Starting from the current actual observation center, the system uses the updated parameter set as the advection driving force and performs a preset number of iterations at a preset time step (e.g., 10 minutes). The logic for calculating the center position in the nth iteration is as follows: First, the updated transmission velocity scalar is multiplied by the updated diffusion dominant direction vector to obtain the driving velocity vector at the current moment. Then, this driving velocity vector is multiplied by the preset time step to obtain the displacement vector for this iteration. Finally, this displacement vector is accumulated to the center position coordinates obtained in the (n-1)th iteration (or, for the first iteration, to the starting point), thereby calculating the center position coordinates for the nth iteration. The generated coordinate sequence is the dynamic pollution diffusion path.

[0075] It is worth noting that the preset time step is determined based on the combined constraints of the Courant-Friedrichs-Lewy (CFL) Condition and the grid spatial resolution limit. The minimum grid edge length of the unified spatiotemporal grid dataset is obtained, and the historical maximum transmission wind speed of the monitoring area is statistically analyzed. A specific computational logic is used to determine the critical value of the time step. The minimum grid edge length of the unified spatiotemporal grid dataset is multiplied by a preset safety factor (e.g., 0.5) to obtain the safe grid span. This safe grid span is then divided by the historical maximum transmission wind speed of the monitoring area, and the quotient is used as the maximum allowable critical value of the time step. This setting aims to ensure that within a time step, the displacement of the virtual particle does not cross more than one complete grid cell, thereby avoiding numerical dispersion and trajectory truncation errors caused by excessively large step sizes, and ensuring the continuity and physical realism of the inference path.

[0076] For example, assuming the coordinates of the pollution center at the previous moment were x-coordinate 100 and y-coordinate 200, the transmission velocity obtained from step S14 was 10 meters per second, the dominant diffusion direction was a horizontal eastward unit vector (i.e., the horizontal component was 1, and the vertical component was 0), and the data acquisition time interval was 60 seconds. When making theoretical predictions, the theoretical x-coordinate at the current moment is calculated by adding the previous moment's x-coordinate value of 100 to the product of the velocity value of 10, the horizontal component value of 1, and the time interval value of 60, resulting in 700. Since the y-coordinate remains unchanged due to the absence of a velocity component, the theoretical prediction point coordinates are (700, 200). For actual observations, the actual high-concentration center coordinates calculated using the centroid method are (760, 200). The difference between theory and reality is calculated, yielding a deviation vector of (60, 0), meaning a horizontal deviation of 60 and a vertical deviation of 0, with a Euclidean distance of 60 meters between them. During parameter correction, the feedback gain coefficient for velocity correction is set to 0.5. The velocity correction is calculated by multiplying the feedback gain coefficient of 0.5 by the distance error of 60, and then dividing by the time interval of 60, resulting in 0.5 meters per second. This correction is added to the original velocity, yielding an updated transmission velocity of 10.5 meters per second. Since there is no longitudinal deviation, the dominant diffusion direction remains unchanged. Finally, path generation is performed. Using the current actual observation center (760, 200) as the starting point, the path for the next hour is extrapolated eastward using the updated velocity of 10.5 meters per second. A path point is recorded every 10 minutes, generating a path sequence containing a series of coordinate points. For example, the sequence includes the starting coordinates (760, 200) and the coordinates of the second point after 10 minutes of extrapolation. .

[0077] In step S16, a preset concentration decay function is applied along the dynamic pollution diffusion path to perform spatial gradient extrapolation, calculate the spatiotemporal concentration distribution at future times, and obtain the final pollution trend prediction result, including:

[0078] Determine the path points on the dynamic pollution diffusion path, take the path points as diffusion centers, and apply the Gaussian distribution function as the concentration decay function.

[0079] The displacement distance of the diffusion center on the time axis is determined using the transmission speed, and the spatial coverage range and amplitude attenuation rate of the Gaussian distribution function are determined using the source strength attenuation coefficient and the preset diffusion coefficient.

[0080] The offset distance of each point in the spatial grid relative to the path center is calculated, and the predicted concentration value of each grid point is calculated by substituting it into the concentration decay function. This generates a spatiotemporal concentration distribution map for future times, which serves as the final prediction result of the pollution trend.

[0081] It should be noted that determining the path points on the dynamic pollution diffusion path, using these path points as diffusion centers, and applying the Gaussian distribution function as the concentration decay function is achieved using a discretized form of the two-dimensional Gaussian Puff Model. For any predicted time on the dynamic path, the system defines the path point at that time as the instantaneous diffusion center. The specific calculation logic of the concentration decay function is as follows: First, for any grid point to be predicted in the spatial grid, calculate the square of the Euclidean distance between it and the path point that is the diffusion center at the current moment. Specifically, calculate the sum of the squares of the differences in their horizontal and vertical coordinates. Second, obtain the diffusion standard deviation used to characterize the spatial coverage of pollutants at the current moment, and calculate twice the square of the diffusion standard deviation. Next, divide the square of the Euclidean distance by twice the square of the diffusion standard deviation, and take the opposite of the quotient (i.e., multiply by negative one). Subsequently, calculate the value of the natural exponential function with the natural constant e as the base and the opposite of e as the exponent. Finally, multiply the current central peak concentration by the value of the natural exponential function, and the product is the predicted concentration of the grid point to be predicted at the current moment.

[0082] It should be noted that determining the displacement distance of the diffusion center on the time axis using the transmission speed refers to directly calling the coordinate sequence in the dynamic pollution diffusion path generated in step S15. Determining the spatial coverage and amplitude decay rate of the Gaussian distribution function using the source strength attenuation coefficient and the preset diffusion coefficient involves calculating the evolution of the center peak concentration and diffusion standard deviation through the physical attenuation equation. Specifically, using the source strength attenuation coefficient obtained from step S14, the amplitude decay rate is calculated using exponential decay calculation logic. First, the product of the source strength attenuation coefficient and the predicted elapsed time is calculated. The negative of this product is taken as the exponent, and the natural exponential function value with the natural constant as the base is calculated. Finally, this function value is multiplied by the center concentration observed at the current time to obtain the center peak concentration at the predicted time. Using the preset diffusion coefficient, the trend of spatial coverage expanding over time is calculated using linear growth calculation logic. Specifically, the product of the preset diffusion coefficient and the predicted elapsed time is calculated, and this product is added to the initial diffusion range value to obtain the diffusion standard deviation at the predicted time.

[0083] It is worth noting that the preset diffusion coefficient is determined based on the Pasquill Stability Classes statistics of the historical meteorological conditions in the monitoring area. Tracer diffusion experimental data or high-precision simulation data were collected from the past three years for this area under different atmospheric stability levels (e.g., extremely unstable Class A to extremely stable Class F). Regression analysis was performed on the growth rate of the plume width over time for each stability class, and the slope of the regression curve was extracted as the empirical diffusion coefficient under that stability condition. During system operation, the current stability class is matched with real-time meteorological data, and the corresponding value is automatically loaded. The value is used as a preset parameter.

[0084] It should be noted that calculating the offset distance of each point in the spatial grid relative to the path center and generating the spatiotemporal concentration distribution map is achieved through grid traversal and matrix operations. For each time step on the prediction path... (For example, at the 10th minute in the future), the system initializes a zero matrix with the same size as the unified spatiotemporal grid dataset. Then, it iterates through each spatial grid point corresponding to this matrix, calculating the Euclidean distance between that grid point and the current path center point. Specifically, this involves calculating the sum of the squares of the differences in their horizontal and vertical coordinates, and then taking the square root of this sum. Next, this Euclidean distance value, along with the calculated central peak concentration and diffusion standard deviation at the current moment, is substituted into the aforementioned Gaussian distribution concentration decay calculation logic to obtain the predicted concentration value for that grid point, which is then filled into the corresponding position in the matrix. Finally, the concentration matrices corresponding to all predicted time steps are stacked and stitched together in chronological order to construct a digital tensor composed of three dimensions: time, spatial height, and spatial width. This is the spatiotemporal concentration distribution map for the future moment.

[0085] For example, assuming the initial state is as follows: the currently observed central concentration is 100 mg / m³, and the initial diffusion range is 50 m. Step S14 analyzes a source strength attenuation coefficient of 0.05 per minute, and the preset diffusion coefficient matched to the current weather conditions is 10. The state is predicted for the next 4 minutes. The central peak concentration, after attenuation calculation, is approximately 81.87 mg / m³. The specific calculation logic is to multiply the initial concentration by a natural exponential function with a base of the natural constant and an exponent of the negative product of the source strength attenuation coefficient and time. The diffusion range is expanded to 70 m. For a grid point 70 m from the center (equivalent to one standard deviation of diffusion), its predicted concentration is approximately 49.65 mg / m³. The specific calculation logic is to multiply the predicted central peak concentration by a natural exponential function with a base of the natural constant and an exponent of -0.5. This value is filled into the corresponding coordinates of the distribution map, forming the final trend prediction data.

[0086] In summary, this invention collects multi-source heterogeneous environmental data and constructs a unified spatiotemporal grid dataset. It utilizes a multi-scale dilated convolutional network to extract spatiotemporal correlation features under different receptive fields and innovatively introduces environmental dynamic change indicators for attention-weighted fusion, generating a comprehensive environmental state vector that combines micro-fluctuations and macro-trends. Furthermore, this invention uses a parametric regression mapping model to analyze deep abstract features into diffusion dynamic parameters with clear physical meaning. It also performs closed-loop feedback calibration of the model through trajectory deviations of real-time observation data, driving the diffusion evolution model and concentration decay function to perform spatiotemporal extrapolation. This effectively solves the technical problems of low data utilization, delayed physical model response, and lack of interpretability in purely data-driven models in traditional monitoring methods, achieving accurate source tracing of pollutant diffusion paths and high-precision prediction of future spatiotemporal concentration distribution under complex dynamic environments.

[0087] Reference Figure 2 This invention also provides a system for processing multi-source heterogeneous environmental monitoring data, comprising:

[0088] The data preprocessing module is used to collect multi-source heterogeneous environmental data, perform protocol parsing on the multi-source heterogeneous environmental data, perform spatiotemporal grid alignment, construct a unified spatiotemporal grid dataset, and calculate environmental dynamic change indicators that characterize data fluctuation characteristics based on the unified spatiotemporal grid dataset.

[0089] The multi-scale feature extraction module is used to perform multi-scale convolution operations on the unified spatiotemporal grid dataset, and use convolution kernels of different sizes to extract the spatiotemporal correlation features of each grid unit under different receptive fields to generate multi-layer spatiotemporal feature maps.

[0090] The state fusion coding module is used to fuse and encode the environmental dynamic change index with the multi-layer spatiotemporal feature map by weighted superposition to generate a comprehensive environmental state vector containing real-time dynamic environmental information.

[0091] The physical parameter inversion module is used to input the comprehensive environmental state vector into a preset parameter regression mapping model and analyze it to obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion.

[0092] The dynamic path correction module is used to deduce the theoretical prediction state at the current moment using the diffusion dynamics parameter set, calculate the state deviation between the theoretical prediction state and the current moment observation value in the unified spatiotemporal grid dataset, calibrate and update the diffusion dynamics parameter set based on the state deviation, and use the updated parameter set to drive the preset diffusion evolution model to generate a future-oriented dynamic pollution diffusion path.

[0093] The trend prediction and extrapolation module is used to perform spatial gradient extrapolation along the dynamic pollution diffusion path using a preset concentration decay function, calculate the spatiotemporal concentration distribution at future times, and obtain the final pollution trend prediction result.

[0094] It should be noted that the multi-source heterogeneous environmental monitoring data processing system provided in this embodiment of the invention is used to execute all the process steps of the multi-source heterogeneous environmental monitoring data processing method of the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0095] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a processing program for multi-source heterogeneous environmental monitoring data. When the processor executes the computer program, it implements the steps described in the various embodiments of the multi-source heterogeneous environmental monitoring data processing method described above, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as the data preprocessing module.

[0096] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0097] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0098] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.

[0099] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0100] If the modules / units integrated into the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or system capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0101] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0102] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for processing multi-source heterogeneous environmental monitoring data, characterized in that, include: Collect multi-source heterogeneous environmental data, perform protocol parsing on the multi-source heterogeneous environmental data, and perform spatiotemporal grid alignment to construct a unified spatiotemporal grid dataset. Based on the unified spatiotemporal grid dataset, calculate environmental dynamic change indicators that characterize data fluctuation characteristics. Multi-scale convolution operations are performed on the unified spatiotemporal grid dataset, and spatiotemporal correlation features of each grid unit under different receptive fields are extracted using convolution kernels of different sizes to generate multi-layer spatiotemporal feature maps. By using weighted superposition, the environmental dynamic change indicators are fused and encoded with the multi-layer spatiotemporal feature map to generate a comprehensive environmental state vector containing real-time dynamic environmental information. The integrated environmental state vector is input into a preset parameter regression mapping model to obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion. The set of diffusion dynamic parameters includes the dominant diffusion direction, transport velocity, and source strength attenuation coefficient. Based on the diffusion dynamics parameter set, the current linear motion equation is constructed, the theoretical predicted coordinates of the pollution center at the current moment are calculated, and the Euclidean distance between the theoretical predicted coordinates and the coordinates of the high concentration center actually observed in the unified spatiotemporal grid dataset is calculated as the state deviation. Using a preset feedback gain coefficient, the transmission velocity and diffusion dominance direction are corrected in reverse according to the state deviation to obtain an updated diffusion dynamics parameter set. The updated parameter set is substituted into a preset Lagrange particle tracking model to iteratively calculate the center position coordinates at future moments until a preset prediction duration threshold is reached, thereby generating a dynamic pollution diffusion path. The spatial gradient is extrapolated by applying a preset concentration decay function along the dynamic pollution diffusion path to calculate the spatiotemporal concentration distribution at future times, and the final pollution trend prediction result is obtained.

2. The method for processing multi-source heterogeneous environmental monitoring data according to claim 1, characterized in that, The process involves collecting multi-source heterogeneous environmental data, parsing the data using protocols, aligning the spatiotemporal grids to construct a unified spatiotemporal grid dataset, and calculating environmental dynamic change indicators characterizing data fluctuations based on the unified spatiotemporal grid dataset, including: Based on the preset protocol parsing rules, fields are extracted from the raw data packets collected from devices from different manufacturers to obtain standard format data containing latitude and longitude, timestamps and monitoring values; The standard format data is mapped to a spatiotemporal grid coordinate system of a preset resolution using the inverse distance weighted interpolation method. The grid cells of missing data are interpolated and filled to obtain a unified spatiotemporal grid dataset. The variance of the monitoring data within each grid cell is calculated based on a time sliding window of a preset length, and the variance is used as an indicator of environmental dynamic change that characterizes the fluctuation characteristics of the data.

3. The method for processing multi-source heterogeneous environmental monitoring data according to claim 1, characterized in that, The process of performing multi-scale convolution operations on the unified spatiotemporal grid dataset, using convolution kernels of different sizes to extract the spatiotemporal correlation features of each grid unit under different receptive fields, and generating multi-layer spatiotemporal feature maps includes: Construct a feature extraction network containing multiple parallel convolutional branches, where each branch is configured with dilated convolutional kernels with different dilation rates; The unified spatiotemporal grid dataset is simultaneously input into the multiple parallel convolutional branches to extract local feature maps at different receptive field scales. The local feature maps output from each branch are concatenated along the channel dimension, and a 1×1 convolutional layer is used for channel dimensionality reduction and feature integration to generate the multi-layer spatiotemporal feature map.

4. The method for processing multi-source heterogeneous environmental monitoring data according to claim 1, characterized in that, The step involves weighted superposition to fuse and encode the environmental dynamic change indicators with the multi-layer spatiotemporal feature map, generating a comprehensive environmental state vector containing real-time dynamic environmental information, including: The environmental dynamic change index is broadcast and expanded in the spatial dimension so that its size is consistent with the spatial resolution of the multi-layer spatiotemporal feature map, thus obtaining the fluctuation feature matrix. A preset feature weight matrix is ​​introduced, and the fluctuation feature matrix is ​​weighted element by element. The weighted fluctuation feature matrix and the weighted multi-layer spatiotemporal feature map are added element by element to obtain the fused feature tensor. The feature tensor is then subjected to global average pooling to generate the comprehensive environmental state vector.

5. The method for processing multi-source heterogeneous environmental monitoring data according to claim 1, characterized in that, The step of inputting the comprehensive environmental state vector into a preset parametric regression mapping model to analyze and obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion includes: The comprehensive environmental state vector is input into the preset parameter regression mapping model and processed by the nonlinear activation function of the hidden layer to be mapped to the physical parameter space; The normalized direction vector is calculated using the Tanh activation function of the output layer as the dominant diffusion direction. The scalar value of the transmission speed and the source strength attenuation coefficient are calculated using the Sigmoid activation function and a preset physical extreme value range.

6. The method for processing multi-source heterogeneous environmental monitoring data according to claim 1, characterized in that, The process involves applying a preset concentration decay function along the dynamic pollution diffusion path to perform spatial gradient extrapolation, calculating the spatiotemporal concentration distribution at future times, and obtaining the final pollution trend prediction result, including: Determine the path points on the dynamic pollution diffusion path, take the path points as diffusion centers, and apply the Gaussian distribution function as the concentration decay function. The displacement distance of the diffusion center on the time axis is determined using the transmission speed, and the spatial coverage range and amplitude attenuation rate of the Gaussian distribution function are determined using the source strength attenuation coefficient and the preset diffusion coefficient. Calculate the offset distance of each point in the spatial grid relative to the path center, substitute it into the concentration decay function to calculate the predicted concentration value of each grid point, and generate a spatiotemporal concentration distribution map for future times.

7. A system for processing multi-source heterogeneous environmental monitoring data, characterized in that, For implementing the method as described in any one of claims 1-6, comprising: The data preprocessing module is used to collect multi-source heterogeneous environmental data, perform protocol parsing on the multi-source heterogeneous environmental data, perform spatiotemporal grid alignment, construct a unified spatiotemporal grid dataset, and calculate environmental dynamic change indicators that characterize data fluctuation characteristics based on the unified spatiotemporal grid dataset. The multi-scale feature extraction module is used to perform multi-scale convolution operations on the unified spatiotemporal grid dataset, and use convolution kernels of different sizes to extract the spatiotemporal correlation features of each grid unit under different receptive fields to generate multi-layer spatiotemporal feature maps. The state fusion coding module is used to fuse and encode the environmental dynamic change index with the multi-layer spatiotemporal feature map by weighted superposition to generate a comprehensive environmental state vector containing real-time dynamic environmental information. The physical parameter inversion module is used to input the comprehensive environmental state vector into a preset parameter regression mapping model and analyze it to obtain a set of diffusion dynamic parameters characterizing the physical properties of pollutant diffusion. The dynamic path correction module is used to deduce the theoretical prediction state at the current moment using the diffusion dynamics parameter set, calculate the state deviation between the theoretical prediction state and the current moment observation value in the unified spatiotemporal grid dataset, calibrate and update the diffusion dynamics parameter set based on the state deviation, and use the updated parameter set to drive the preset diffusion evolution model to generate a future-oriented dynamic pollution diffusion path. The trend prediction and extrapolation module is used to perform spatial gradient extrapolation along the dynamic pollution diffusion path using a preset concentration decay function, calculate the spatiotemporal concentration distribution at future times, and obtain the final pollution trend prediction result.

8. A processing device for multi-source heterogeneous environmental monitoring data, characterized in that, include: A processor and a memory connected to the processor; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to cause the processor to perform the method for processing multi-source heterogeneous environmental monitoring data as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, the computer instructions being configured to cause the computer to perform a method for processing multi-source heterogeneous environmental monitoring data according to any one of claims 1-6.

Citation Information

Patent Citations

  • Ecological meteorology and satellite remote sensing combined environment dynamic monitoring method and system

    CN120510942A

  • Water environment pollution tracing system and method

    CN121072283A