Intelligent environmental protection dynamic monitoring system based on multi-source data fusion
By using a smart environmental dynamic monitoring system that integrates multi-source data and employs Markov iterative prediction and conditional risk value indicators, a multi-agent collaborative action space is constructed. This solves the collaborative decision-making problem in scenarios of sudden pollution spread and multi-agent collaborative emission reduction in existing technologies, and achieves proactive early warning and efficient collaborative control.
Patent Information
- Application Number
- CN202610809758.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-25
AI Technical Summary
Existing intelligent environmental protection dynamic monitoring systems lack a consistent collaborative decision-making mechanism based on full-state Markov evolution when dealing with sudden pollution spread and large-scale multi-entity collaborative emission reduction scenarios. This leads to distortion of local risk assessments due to the nonlinear evolution trajectory of pollution sources, inefficient sharing of monitoring information, delayed response to environmental governance, and serious waste of resources.
The intelligent environmental protection dynamic monitoring system, which integrates multi-source data, includes modules for data acquisition, status assessment, time series prediction, risk quantification, and collaborative control. It utilizes Markov iterative prediction and conditional risk value index to construct a multi-agent collaborative action space, thereby achieving real-time collaborative control.
It enhances the proactive early warning capability for pollution risks, ensures the safety and efficiency of control strategies, reduces communication overhead, and achieves efficient collaborative control in spatially heterogeneous scenarios.
Smart Images

Figure CN122635933A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental monitoring technology, specifically to a smart environmental dynamic monitoring system based on multi-source data fusion. Background Technology
[0002] The engineering implementation of intelligent environmental dynamic monitoring systems is a core evolutionary direction in the current fields of environmental governance and the Industrial Internet of Things. It aims to construct a risk early warning and collaborative control system with environmental adaptability through multi-source sensing data fusion and spatiotemporal evolution modeling. In recent years, with the development of high-precision industrial sensors and distributed computing technology, the research and development paradigm of environmental monitoring systems is rapidly evolving towards real-time operation, multimodal fusion, and intelligent decision-making collaboration.
[0003] Existing technologies typically employ an architecture based on classical predictive control algorithms combined with static threshold triggering to achieve pollution status monitoring and automatic spray emission reduction on regional environmental monitoring platforms. These solutions can achieve basic data coverage when handling routine industrial emission scenarios and can perform a certain degree of pollution trend assessment using historical statistical data, possessing some regional monitoring capabilities. However, for scenarios involving sudden pollution spread and large-scale multi-entity collaborative emission reduction, the pollution evolution prediction model and risk assessment mechanism are severely logically decoupled. Each monitoring unit responds independently based on local observations, lacking a consistent collaborative decision-making mechanism based on full-state Markov evolution.
[0004] Because environmental pollution exhibits high non-stationarity and strong spatiotemporal correlation during dynamic diffusion, and environmental governance facilities inherently suffer from execution lags in response scheduling and multi-objective coordination, this decision-making paradigm based on isolated units and single mean predictions is highly susceptible to failure. When the system encounters sudden pollution diffusion under extreme weather conditions or high-density, multi-site coordinated emission reduction demands, the nonlinear evolution trajectory of pollution sources can lead to instantaneous distortion of local risk assessments. Inefficient sharing of monitoring information among different actuators and decision-making blind spots cause severe control deviations, resulting in significant systemic lags or even local loss of control in environmental governance responses. This leads to a substantial waste of environmental protection resources and severe impairment of monitoring and early warning effectiveness. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a smart environmental dynamic monitoring system based on multi-source data fusion.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] The data acquisition module is used to acquire multi-source environmental raw data of the target area and perform spatiotemporal synchronization and alignment to generate a multi-dimensional environmental feature tensor.
[0008] The state assessment module is used to extract the concentration feature values of the multidimensional environmental feature tensor and map them to a preset pollution level range, and output the pollution discrete state vector at the current moment.
[0009] The time-series prediction module is used to call the transition probability matrix obtained by training from historical environmental samples, and perform Markov iteration prediction on the pollution discrete state vector based on the transition probability matrix to generate a pollution probability distribution sequence within a preset decision period.
[0010] The risk quantification module is used to extract geographic coordinate features from the multidimensional environmental feature tensor, perform spatiotemporal weighting by combining the expected values at each time point in the pollution probability distribution sequence, and generate a dynamic risk field function; the dynamic risk field function is then pre-set with a confidence level of... The tail integral operation generates the conditional value at risk index;
[0011] The collaborative control module is used to receive the real-time state parameters of each actuator corresponding to the multi-dimensional environmental feature tensor, and construct a multi-agent collaborative action space by combining the geographical location weights of each actuator; the pollution discrete state vector, the dynamic risk field function and the conditional risk value index are synchronously input into a preset risk constraint reinforcement learning model, and the optimal collaborative control strategy is calculated based on the multi-agent collaborative action space under the constraint that the conditional risk value index is lower than a preset risk threshold.
[0012] The instruction execution module is used to output the optimal cooperative control strategy, receive environmental feedback data after execution, and send the environmental feedback data back to the data acquisition module to update the multidimensional environmental feature tensor.
[0013] This invention discloses a smart environmental dynamic monitoring system based on multi-source data fusion, comprising:
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0015] 1. This invention integrates Markov prediction results with spatial geographic features, enabling the system to assess the tail risk level before pollution occurs, thereby upgrading the passive response alarm mechanism into a proactive predictive risk warning mechanism.
[0016] 2. This invention introduces the Conditional Value at Risk (CVaR) index as the core risk constraint target into the optimization process, and ensures the safety of the control strategy through a two-layer protection mechanism: at the optimization algorithm level, it is used as a penalty constraint; at the execution decision level, a final safety defense line that can directly intervene in the action selection is set.
[0017] 3. This invention encodes geographical location weights into the action space construction process, enabling spatially adjacent actuators to form implicit collaborative relationships, while actuators that are far apart remain relatively independent. This reduces communication overhead and achieves efficient collaborative control in spatially heterogeneous scenarios. Attached Figure Description
[0018] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:
[0019] Figure 1 This is a system module connection diagram of the present invention;
[0020] Figure 2 This is a flowchart illustrating the working sequence of the present invention. Detailed Implementation
[0021] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.
[0022] In existing technologies, intelligent environmental dynamic monitoring systems mostly rely on single-point threshold exceedance alarms or static rule control, lacking probabilistic modeling of the temporal evolution trend of pollution states and quantitative constraint mechanisms for extreme risks. The data processing layer of existing solutions does not systematically address the spatiotemporal asynchrony of multi-source sensors, resulting in coordinate offsets and timestamp misalignments in the fused data, thus affecting the accuracy of state estimation. At the control layer, existing solutions generally lack modeling of the geographical distribution and response differences of actuators, failing to maximize the synergistic gain among multiple actuators. Regarding adaptive mechanisms, existing solutions lack quantitative criteria for automatically triggering retraining after model drift occurs, leading to a degradation in prediction accuracy over time.
[0023] To address the aforementioned issues, this invention reveals that the pollution diffusion process exhibits Markov property in the time dimension, meaning that the probability distribution of the pollution state at the next moment depends only on the current state vector and external meteorological driving factors, and is independent of historical state conditions at earlier moments. In the spatial dimension, pollutant migration is significantly influenced by topographic damping effects. Establishing a dynamic risk field function with physical diffusion semantics can effectively characterize the spatial non-uniformity of pollution diffusion. Further research shows that introducing Conditional Value at Risk (CVaR) as a hard constraint into the reinforcement learning loss function can ensure optimal control performance while avoiding extreme pollution events, achieving optimal collaborative control with manageable risk.
[0024] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0025] Example:
[0026] like Figure 1 As shown, the intelligent environmental protection dynamic monitoring system based on multi-source data fusion includes: a data acquisition module, a status assessment module, a time series prediction module, a risk quantification module, a collaborative control module, an instruction execution module, and a self-evolution correction module.
[0027] like Figure 2 The diagram shown is a flowchart illustrating the operational sequence of this invention. The data acquisition module is deployed at the edge computing node of an industrial park or urban environmental monitoring center. It collects raw measurement data in real time from various types of sensors within the target area via a dual-link system: wired industrial Ethernet (such as ModBus / TCP protocol) and wireless Internet of Things (such as NB-IoT or LoRa protocol). This data includes, but is not limited to, gaseous pollutant concentration sensors at fixed sites (measuring SO2, NO...). x The data streams include data from sensors measuring ozone, PM2.5, etc. (sampling period of 1 minute), meteorological sensors (measuring wind speed, wind direction, temperature, and humidity, sampling period of 10 seconds), and online continuous emission monitoring systems (CEMS, sampling period of 5 seconds) at each emission source. Each raw data record includes the sensor device number, UTC timestamp (accurate to milliseconds), and WGS-84 coordinates (latitude, longitude, and altitude). In other embodiments, the raw measurement data may also include water quality sensor data, noise monitoring data, video image data, etc.
[0028] The specific implementation steps for spatiotemporal synchronization alignment are as follows:
[0029] 1. Using the unified time base set by the system (such as whole minute UTC time) as the anchor point, perform linear interpolation on each asynchronous data stream to uniformly resample data with different sampling frequencies to a fixed time step Δt (e.g., Δt=1 minute).
[0030] 2. Map the geographic coordinates of all sensors to a preset latitude and longitude spatial grid, with the grid resolution set to Δx×Δy (e.g., 100m×100m). Take the area-weighted average of the measurement values of multiple sensors within the same grid.
[0031] 3. Organize the aligned multi-source data according to the structure of [number of spatial grid rows × number of spatial grid columns × number of time steps × number of feature channels] to generate a multidimensional environmental feature tensor. Where M and N are the number of rows and columns of the spatial grid, K is the number of time steps, and F is the number of feature channels (including dimensions such as concentration, meteorology, and topography).
[0032] The state assessment module uses a weighted average or Gaussian process regression method to fuse representative concentration values from each grid channel of the multidimensional environmental feature tensor T, extracting a subset of concentration feature values, i.e., the measured concentration values of each pollutant in each grid cell at the current time t. (i=1,2,...,n, where n is the number of pollutant types). The concentration values of each pollutant are classified according to a preset pollution level mapping table (set with reference to the National Ambient Air Quality Standard GB3095-2012 or local emission standards, dividing the concentration range into L discrete levels, for example, L=6, corresponding to excellent, good, lightly polluted, moderately polluted, heavily polluted, and severely polluted), generating discrete level codes for each pollutant. ∈{1,2,...,L}. The discrete level codes of n pollutants are concatenated to generate the discrete pollution state vector at the current moment. ∈ This vector represents the overall pollution status of the target area at the current moment using concise discrete notation.
[0033] During the system initialization phase, the time series prediction module obtains the state transition probability matrix offline through maximum likelihood estimation using historical environmental samples (containing at least three years of continuous monitoring data). Matrix elements This represents the probability of transitioning from state i to state j within one time step Δt, satisfying the normalization condition. During online runtime, the current discrete state vector of pollution is used. Given the initial distribution, perform k-step iterative prediction using matrix exponentiation: , Let be the pollution probability distribution vector at step k in the decision-making period, satisfying Repeat the above operation for all time steps k=1,2,...,H within the preset decision period H to generate a pollution probability distribution sequence. .
[0034] To address the problem of the state space dimension exploding with the increase of pollutant types, this invention employs one or a combination of the following dimensionality reduction strategies:
[0035] (1) State aggregation: Based on the correlation of pollutant concentration levels, The original states are clustered into L' macro states (L'<< The transition probability matrix is trained in the macro-state space. It is automatically enabled when n≥4.
[0036] (2) Sparse estimation: Utilizing the sparsity of the transition probability matrix, only non-zero transition terms are stored, and the compressed sparse row format is used for storage.
[0037] (3) Tensor decomposition: Utilizing the conditional independence between pollutants, the joint state transition is decomposed into the marginal distribution product of the independent state transitions of each pollutant.
[0038] The risk quantification module receives geographic coordinate features (latitude and longitude coordinates of the center point of each grid cell) from the multidimensional environmental feature tensor T, and combines them with the pollution probability distribution sequence. Expected pollution level values at various times Spatiotemporal weighted calculations are then performed. First, the coordinates of each known pollution source are used... Call the Gaussian kernel function with the center point as the reference point. Spatial diffusion mapping is performed on the expected values at each time step:
[0039] ;
[0040] in , Let (x, y) be the diffusion scale parameter of the s-th pollution source (calibrated by historical diffusion experiments, unit: meters), and (x, y) be the target evaluation grid coordinates. This represents the initial risk distribution value. Subsequently, the topological features of the target area (including the digital elevation model (DEM) and feature shading coefficients) are retrieved from a pre-defined terrain database. The damping coefficient is then calculated based on the terrain slope θ(x,y) and the shading coefficient ρ(x,y). By multiplying the damping coefficient point by point with the initial risk distribution, a dynamic risk field function with physical diffusion semantics is generated:
[0041] ;
[0042] Dynamic risk field function This is a dimensionless indicator, with its value range mapped to [0,1], where 1 represents extremely high risk. A pre-set confidence level α (range α∈[0.9,0.99], typical value α=0.95) is used to... Tail integration is performed on the joint distribution in the spatial and temporal domains to calculate the conditional value-at-risk index:
[0043] ;
[0044] in Value at risk (α quantile of the RF distribution) corresponding to a confidence level α (range α∈[0.9,0.99]). Defined as under the probability distribution of the risk field function, for more than The expected value is obtained by calculating the tail region of the quantile. This single scalar characterizes the expected risk level of extreme pollution events within the decision-making period and is used for constraint determination in subsequent collaborative control modules.
[0045] The collaborative control module receives real-time status parameters from various actuators (such as air purification devices, sprinkler systems, and emission source shut-off valves) corresponding to the multi-dimensional environmental feature tensor T. These real-time status parameters are acquired in real-time via actuator feedback interfaces or industrial control buses, including current on / off status (0 / 1), operating power (kW), and remaining consumables (%). This is combined with the geographical coordinates of each actuator. Calculate the geographical location weight of each actuator relative to the target control area. :
[0046] ;
[0047] Where Ω represents the grid set of the target control area, and |Ω| represents the total number of grids. The effective control radius (in meters) of the actuators is defined. A multi-agent cooperative action space is constructed based on the state parameters and geographical location weights of each actuator. Where m is the total number of actuators, This represents the discrete action set of the a-th actuator (e.g., q+1 levels: off, low-power operation, medium-power operation, and high-power operation). To avoid the influence of absolute weight values, the weights of all actuators will be normalized across actuators before subsequent use. , making , where m is the total number of actuators.
[0048] The risk-constrained reinforcement learning model adopts the Constrained Markov Decision Process (CMDP) framework. Its state input is... The concatenated vectors, the action output is the combined action in the multi-agent cooperative action space A, and the constraints are as follows: (Preset risk threshold, unit consistent with RF units). The model solves by:
[0049] ;
[0050] Calculate the optimal cooperative control strategy π*, where γ∈(0,1) is the discount factor (typically 0.95).
[0051] The instruction execution module transforms the optimal cooperative control strategy π* into specific control instructions for each actuator, which are then sent to the hardware actuators via the industrial control bus. After execution, the updated environmental data collected by each sensor serves as environmental feedback data and is transmitted back to the data acquisition module via the communication link, triggering the rolling update of the multidimensional environmental feature tensor T, thus forming a complete closed-loop control circuit.
[0052] This embodiment employs the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. Each actor's policy network is a 4-layer fully connected network with an input layer dimension of 817, two hidden layers of 400 and 300 nodes respectively, and an output layer mapped to the action space using Tanh activation. The evaluation network takes the global state and the joint actions of all actors as input, with hidden layers of 400 and 300 nodes each, and outputs a single Q-value. Hyperparameter settings are as follows: Adam optimizer, policy network learning rate 1e-4, evaluation network learning rate 1e-3, discount factor γ = 0.95. Experience replay pool capacity 1e6, batch size 1024. Target network soft update coefficient τ = 0.01. Total training steps 1e7, evaluation every 1000 steps. Exploration noise is handled using the Ornstein-Uhlenbeck process with parameters θ = 0.15 and σ = 0.2.
[0053] It should be noted that the risk-constrained reinforcement learning model adopts a two-stage operation mode of "offline centralized training and online distributed inference": In the offline stage, the policy network and evaluation network are pre-trained based on historical environment samples until the model converges; in the online stage, the parameters of the evaluation network are frozen, and only the policy network is updated with a lightweight gradient correction to adapt to environmental fluctuations. The two do not share a time scale, and the triggering frequency of offline training is much lower than the decision frequency of online inference.
[0054] Through the above technical solution, the system, through the coordinated cooperation of six functional modules, synchronizes and aligns multi-source heterogeneous sensor data in both time and space dimensions. After a complete process of state discretization, Markov iterative prediction, spatiotemporal risk field construction, and conditional risk value constraint optimization, it outputs a safe and controllable collaborative control strategy, enabling the environmental monitoring system to transform from a passive post-event response to an active predictive management system.
[0055] In one embodiment, before generating the multidimensional environmental feature tensor T, the data acquisition module performs the following two preprocessing steps.
[0056] (I) Outlier Identification and Removal. For the raw environmental data collected from multiple sensors, an improved Z-score method is used to identify outlier observations dimension by dimension. Let the historical mean of the i-th feature channel within the time window [tW,t] (W is the window width, typically W=60 sampling points) be... Standard deviation is Then the current observation value The Z-score is ;
[0057] when (Outlier threshold, assuming sensor measurement noise follows a Gaussian distribution, range) Typical values When ), determine Outlier observations are marked as missing and removed from the sequence. For cases where multiple consecutive time steps exceed the threshold within a short period (e.g., more than 3 consecutive sampling points), the system issues a sensor anomaly alarm to maintenance personnel and suspends further use of the channel's data until manual verification. After the above processing, a cleaned environmental data sequence is obtained, where valid observations are retained and outlier locations are replaced with missing value markers.
[0058] (ii) Spatiotemporal Completion of Missing Dimensions. For missing dimensions in the cleaned environmental data sequence (including missing points caused by outlier removal and complete missing segments due to sensor communication interruptions), a pre-trained Generative Adversarial Network (GAN) is invoked to perform spatiotemporal feature completion. The generator G of the GAN adopts an encoder-decoder structure, taking the complete feature sequence of non-missing dimensions as conditional input and outputting the reconstructed value of the missing dimensions.
[0059] The GAN adopts the WGAN-GP architecture. The generator G is an encoder-decoder structure, with the encoder containing 5 convolutional layers (3×3 kernels, 64→512 channels, stride 2), each followed by batch normalization and LeakyReLU activation; the decoder has 5 deconvolutional layers, with the output layer using Tanh activation. The discriminator D is a 4-layer convolutional network (4×4 kernels, 64→512 channels), each followed by LeakyReLU and Dropout (0.5). Training uses the Adam optimizer with a learning rate of 0.0002. A total of 500 epochs are trained, with a batch size of 32, and the generator is trained twice after each discriminator training epoch. The loss function is Wasserstein distance plus gradient penalty (coefficient 10). The training dataset comes from 12 months of complete historical monitoring records, and training samples are constructed by simulating a missing rate of 10%~40% using random masks.
[0060] During inference, the generator outputs a completion value for each missing position, concatenates the completed data with the original data of the non-missing dimensions to generate a multi-source environmental data matrix with complete dimensions, and then constructs a multi-dimensional environmental feature tensor T according to the aforementioned process.
[0061] The above technical solution introduces two preprocessing steps—improved Z-score outlier removal and generative adversarial network spatiotemporal completion—before constructing the multidimensional environmental feature tensor. This eliminates data quality defects caused by sensor failures, communication interruptions, and asynchronous spatiotemporal sampling, thereby improving the effective data coverage of the input tensor. It ensures the accuracy of subsequent state assessment and time series prediction from the data source and solves the inherent defects of existing solutions, such as coordinate offset and timestamp misalignment in fused data caused by the spatiotemporal asynchrony of multiple sensor sources.
[0062] Based on the standard Markov model prediction process, the time series prediction module superimposes an online dynamic correction mechanism based on the joint change rate of meteorology and emission intensity to enhance the response ability of the transition probability matrix to real-time environmental changes.
[0063] The specific implementation steps are as follows: Extract the non-pollution dimension features from the multi-dimensional environmental feature tensor T, including: wind speed (m / s) and relative humidity (%) at the current moment. At the same time, extract the real-time monitored emission intensity (g / s) from the online monitoring data of each emission source. Before entering the calculation, the above three types of data are normalized to the interval [0, 1] (divided by their respective historical maximum values ).
[0064] Calculate the joint change rate of wind speed, humidity and emission intensity :
[0065] ;
[0066] Where , , are the influence weight coefficients of each factor ( , empirically taken as , , ). The system automatically truncates the value range of to [-1, 1]. A positive value indicates that the diffusion condition deteriorates (the trend of pollutant accumulation increases), and a negative value indicates that the diffusion condition improves (the trend of pollutant dissipation increases).
[0067] According to , perform online reweighting on the state transition weights in the transition probability matrix P. The reweighting rule is: for the transition with an increase in the state level (i.e., the transition with increasing pollution, j > i), the corrected weight is ; for the transition with a decrease in the state level (i.e., the transition with decreasing pollution, j < i), the corrected weight is ; where δ is the modulation coefficient (the value range is 0 < δ ≤ 1, and the typical value is δ = 0.3). After correction, renormalize each row to ensure , and generate the dynamically corrected transition probability matrix P'. Replace P with P' for subsequent iterative prediction, so that the prediction results can respond to the changes in meteorology and emission intensity in real time, and improve the short-term prediction accuracy.
[0068] By using the above technical solution, the joint rate of change of wind speed, humidity and emission intensity is mapped in real time to the reweighting coefficients of each row of the transition probability matrix. This enables the prediction model to dynamically adapt to meteorological changes and emission fluctuations without retraining. The accuracy of pollution state prediction when meteorological conditions change abruptly is improved compared to the static transition matrix solution, and the prediction drift phenomenon of the model in non-stationary environments is effectively suppressed.
[0069] In one embodiment, the risk quantification module generates a dynamic risk field function RF(x,y,t+k), which uses the coordinates of each pollution source in the geographic coordinate features. Centered on (s=1,2,...,S, where S is the total number of pollution sources), the probability distribution sequence of pollution is analyzed. Expected pollution level at step k A preliminary risk distribution is generated by performing spatial diffusion mapping using a Gaussian kernel function.
[0070] Among the diffusion scale parameters Calibrated by historical diffusion experiments or CFD simulations, the unit is meters, and the value ranges from 50 to 500 meters. The value reflects the contribution of the s-th pollution source to the pollution risk at coordinates (x,y) after k steps, without considering terrain obstacles.
[0071] The topological features of the target area are retrieved from the preset terrain database (stored as a raster DEM file with the same resolution as the spatial grid). The slope θ(x,y) (°) and the land cover shading coefficient ρ(x,y) (value range [0,1], 1 indicates complete shading) of each grid cell (x,y) are extracted. The damping coefficient D(x,y) is calculated using the following formula:
[0072] ;
[0073] Where β∈[0.5,2.0] is the shading attenuation coefficient (calibrated by measured diffusion experiments), κ∈[0,1] is the slope damping weight, and β and γ are obtained by fitting historical diffusion experimental data using the least squares method. Take 45°. D(x,y)∈[0,1], the smaller the value, the stronger the terrain's obstruction of pollution diffusion.
[0074] Preliminary risk distribution With damping coefficient Multiplying point by point:
[0075] ;
[0076] in, This is the preliminary revised risk distribution.
[0077] Subsequently, based on the real-time wind speed vector (u,v) and the turbulent diffusion coefficient, the atmospheric advection-diffusion equation was solved. Perform spatiotemporal transport evolution calculations:
[0078] ;
[0079] in, Let S be the horizontal turbulent diffusion coefficient, and S be the linear loss term.
[0080] Solving the above equations yields the risk distribution after spatiotemporal evolution. As the final dynamic risk field output, It includes both time-series predictions of pollution states and topographic physical constraints, possesses clear diffusion semantics, and can be directly used for subsequent CVaR calculations and collaborative control optimization.
[0081] The above technical solution, by combining the Gaussian kernel spatial diffusion map with the point-by-point product of the terrain damping coefficient, generates a dynamic risk field function. It incorporates both temporal probability information on pollution diffusion and topographic physical constraints, making the representation accuracy of risk spatial distribution superior to planar diffusion models that ignore the influence of terrain. In complex terrain scenarios with mountainous cover or dense building clusters, the risk assessment error can be reduced by more than 40%, thus providing a reliable basis for the precise spatial positioning of subsequent control strategies.
[0082] Furthermore, this application proposes that the collaborative control module incorporates the conditional value of risk index during the training and inference process of the risk constraint reinforcement learning model. As a Lagrange multiplier A loss function is introduced to achieve hard-constraint optimization for extreme pollution events.
[0083] Specifically, the original unconstrained policy optimization objective is transformed into a Lagrange augmented objective function:
[0084] ;
[0085] Where λ≥0 are Lagrange multipliers (i.e., ... (penalty coefficient) A preset risk threshold is set. During the training phase, λ is updated alternately using the dual gradient ascent method and the policy gradient: when > When λ increases (increasing the intensity of punishment for high-risk actions); when ≤ As λ decreases (relaxing constraints and allowing for a larger exploration space), the strategy maximizes long-term returns while satisfying risk constraints. The update step size of λ is set to... (Value range: 0.001~0.01).
[0086] During the inference phase, the collaborative control module monitors in real time. and The difference When ΔV is below the preset fluctuation range (Right now Approaching the risk threshold, with values such as When this happens, the system automatically compresses the exploration boundary of the multi-agent cooperative action space: the set of actions of each actuator... The medium-to-high power action settings are temporarily removed from the selectable set (i.e., compressed from {Off, Low Power, Medium Power, High Power} to {Off, Low Power, Medium Power}), generating a constrained action space A'⊆A. Based on A', action selection is performed to ensure that when extreme risks approach the threshold, the system automatically avoids aggressive control actions that might trigger threshold exceedances, achieving preventative risk management.
[0087] The system's risk constraints are jointly guaranteed through two levels:
[0088] 1. Optimization Layer Constraints (Adaptive Penalty): The Lagrange multiplier method is used to handle CVaR constraints, aiming to guide the policy network to learn optimal behavior under long-term risk budget during the training phase. This method is an effective means of handling constrained optimization problems in stochastic environments, with the goal of asymptotically satisfying constraints with high probability.
[0089] 2. Execution-layer hard constraints (absolute safety boundary): When the CVaR value calculated in real-time during online inference approaches the preset risk threshold, it indicates that the system is in a high-risk critical state. At this time, the absolute safety principle, which takes precedence over the optimization objective, is activated. The system forcibly compresses the action space, removing operation levels such as "high-power actions" that may bring greater uncertainty. This logic is based on the following design considerations: In a risk-critical state, the primary goal is to prevent triggering the threshold, rather than pursuing optimal performance. Although the retained medium- and low-power actions have slower responses, their impact is more predictable, providing a more certain and stable risk reduction path, which conforms to the engineering design principle of "safety first". These two mechanisms together constitute a complete constraint system from "optimization guidance" to "safety fallback".
[0090] By incorporating the Conditional Value at Risk (CVR) index as a Lagrange multiplier into the loss function, hard-constraint optimization of the tail risk of extreme pollution events is achieved. When the system detects that the CVR index is close to a preset threshold, it automatically compresses the exploration boundary of the action space, reducing the default rate beyond the threshold to near zero while ensuring control effectiveness. Compared with the reinforcement learning baseline scheme without constraints, the frequency of extreme pollution events is reduced, achieving optimal collaborative control under controllable risk conditions.
[0091] This application further proposes that when constructing the multi-agent cooperative action space, the cooperative control module adopts a global-local hierarchical partitioning logic, divides the actuators into two sets of agents according to the differences in their response characteristics, generates a coarse-grained emission reduction action subspace and a high-frequency disturbance elimination action subspace respectively, and then generates a complete multi-agent cooperative action space through Cartesian product operation.
[0092] Extract response parameters from the real-time status parameters of each actuator, including: the response delay from the issuance of the command to the device reaching the target operating condition. (seconds), minimum effective control cycle for a single action (seconds). Among these, response latency... The specific measurement method is as follows: A standard state switching command (e.g., switching from "off" to "medium power operation") is sent from the system's host computer to the actuator. Simultaneously, a high-precision timer (accuracy ≤ 10 milliseconds) is started, and the status signal fed back by the actuator (such as a motor current sensor or position feedback switch) is monitored in real time. Timing stops when the feedback signal first stabilizes at the target state (maintained for ≥ 1 second), and this time difference is recorded. The same actuator is measured 10 times repeatedly. After removing the maximum and minimum values, the arithmetic mean is taken as the actuator's response delay. For passive devices whose status cannot be directly sensed (such as stationary nozzles), the status can be indirectly measured through changes in downstream sensors (such as pressure switches or flow meters).
[0093] Set response latency classification threshold (Typical value 60 seconds): Response latency Actuators (such as large purification towers and process control valves, which are characterized by slow response and wide coverage) are classified into the global response intelligent agent set. Response latency Actuators (such as fast-switching solenoid valves and mobile purification vehicles, characterized by fast response and small coverage) are categorized into a set of locally compensated intelligent agents. .
[0094] Extract the macroscopic gradient features of the dynamic risk field function RF(x,y,t+k), i.e., calculate the gradient of RF in the spatial domain ∇RF=(∂RF / ∂x,∂RF / ∂y), identify the top p regions with the largest gradient magnitudes in the risk field (p is determined by the number of global actuators), and input the coordinates and intensity of each gradient peak region as feature vectors into the global response agent set. The executor policy network in the middle outputs a coarse-grained decomposition action subspace. (Including the operating condition adjustment combination of each global actuator).
[0095] Extracting the multidimensional environmental feature tensor T in the recent (Values range from 3 to 10 time steps, with a typical value of 5 time steps, i.e., 5 minutes) Instantaneous fluctuation characteristics: Calculate the short-time variance for each characteristic channel. The coordinates and amplitudes of the Q spatial grids with the highest wave intensity are extracted as feature vectors and input into the set of local compensation agents. The actuator policy network in the middle outputs a high-frequency disturbance elimination action subspace. (Including combinations of fast-response actions from various local actuators).
[0096] To achieve organic synergy between global macro-control and local fine-tuning compensation, a global response agent set is needed. With local compensation agent set A hierarchical conditional strategy interaction mechanism is adopted between them:
[0097] First, the global response agent set Each policy network in the system takes macroscopic gradient features as input and outputs a coarse-grained basic action independently (or jointly). This constitutes a coarse-grained emission reduction action subspace. A specific element within it. This coarse-grained basic action. This represents the overall regulatory tone under the current macroeconomic risk gradient, characterized by slow response and large-scale equipment (such as purification towers operating at reduced power).
[0098] Secondly, the established coarse-grained basic actions By concatenating with instantaneous fluctuation features, an enhanced feature vector is formed, which serves as a set of local compensation agents. Additional inputs to each policy network. In other words, the decision-making of the locally compensated agent. In the case of known macroscopic actions The conditional response made under the premise of [the above] is used for fine-tuning of high-frequency disturbances. Eliminating action subspace from high-frequency disturbances Selected from the options.
[0099] Finally, the complete multi-agent cooperative action space A is implicitly defined through the joint distribution of conditional probabilities: In actual execution, the system first samples according to the global strategy. And then according to this Sample from local policies and the current environment state. The final combination of actions It is then sent to the corresponding actuator hardware.
[0100] As an equivalent alternative, A can also be explicitly constructed as and Cartesian product And combined with the hierarchical reinforcement learning (HIR) framework in reinforcement learning, the top-level policy ( ) and underlying strategy ( End-to-end training is performed using an option structure. This embodiment prioritizes the former conditional sampling scheme because it better aligns with the physical intuition of "macro-level decision-making guiding micro-level compensation" and effectively reduces the complexity of the joint action space. Each element in this space represents a complete collaborative action combination of the global and local actuators, covering a full-scale joint control scheme from macro-level trend control to micro-level fluctuation suppression.
[0101] The above technical solution divides the actuators into two sets of intelligent agents: global response and local compensation, based on response delay. This generates a coarse-grained emission reduction action subspace and a high-frequency disturbance elimination action subspace, respectively. These are then combined into a complete cooperative action space via Cartesian product, achieving full-scale joint control of macro-trend management and micro-fluctuation suppression. Compared with a flat single-layer action space design, this hierarchical structure reduces the effective dimension of the policy search space by about 60% in scenarios with a large number of actuators, significantly improving both training convergence speed and control real-time performance.
[0102] This application further proposes that, in the collaborative control module, the training and inference of the risk-constrained reinforcement learning model are based on the following global reward function:
[0103] ;
[0104] Where t is the discrete time step index (t=0,1,...,T), and T is the preset decision period (unit: number of time steps; the decision period can be set according to monitoring needs, such as 30 minutes, 2 hours, etc.; for example, T=60 corresponds to 1 hour). The definitions of each item are as follows:
[0105] The measure of environmental quality improvement is extracted from environmental feedback data. ,in Let be the surface average concentration (mg / m³) of the i-th pollutant at time t. This is the national secondary concentration limit for this pollutant. The health weight coefficients for each pollutant (set according to the pollutant toxicity level standards, meeting the requirements) ). Map it to the interval [0,1], and... , Maintain consistency. For simplicity, we will continue to use... It refers to the value after normalization transformation.
[0106] Energy consumption items are extracted from the real-time status parameters of each actuator: ,in Let be the actual operating power (kW) of the a-th actuator at time t. The normalized total rated power (kW) of all actuators in the system. ∈[0,1].
[0107] The conditional value at risk (VaR) is the value calculated at the current time. Value (consistent with RF dimensions, normalized to [0,1] by max-min and then substituted into the formula). The normalization is based on records in the historical database. The historical maximum and minimum values are normalized offline.
[0108] and The preset balancing weight coefficients, specifically, The energy consumption penalty weight (ranging from 0.1 to 0.5, with a typical value of λ=0.2) is used to suppress excessive operation of actuators in the range of diminishing marginal returns to emission reduction benefits. The risk penalty weight (ranging from 0.1 to 0.5, with a typical value of η=0.3) is used to ensure that the model is continuously constrained by extreme risk penalties while optimizing emission reduction effects. The setting follows these principles: under normal operating conditions ( far below When ), it can be appropriately reduced. (like =0.1), to give the model greater freedom in emission reduction optimization; when near Automatically lift (like =0.5), strengthen the priority of risk management.
[0109] By introducing environmental quality improvement, energy consumption penalty, and conditional risk value penalty into the reward function, and setting weight coefficients that can be dynamically adjusted according to the risk level, the system can effectively constrain the two types of negative benefits of actuator energy consumption and extreme risks while optimizing emission reduction effects. This overcomes the shortcomings of traditional single-objective optimization methods in environmental control, which often fail to address all aspects, and achieves Pareto improvement under multi-objective joint optimization.
[0110] In one embodiment, the risk-constrained reinforcement learning model adopts a distributed architecture, and the specific execution steps of the collaborative control module include: concatenating the gradient features of the dynamic risk field function RF(x,y,t+k) in the spatial and temporal domains (including ∂RF / ∂x, ∂RF / ∂y and the temporal difference ∂RF / ∂t of each grid cell) into a global state feature vector. As a global state variable, it is sent to a preset parameter server via a high-speed local area network. The parameter server aggregates environmental feedback data from all actuators in the entire domain and calculates the global context feature tensor. (m is the number of executors, d' is the local feature dimension of each executor), and broadcasts it back to the local processing unit of each executor.
[0111] Each executor has its own independently deployed value network. global context feature tensor Using the set of candidate actions in the current multi-agent cooperative action space A as joint input, the actuator is calculated to evaluate the effect of the candidate actions on the target. Long-term cumulative impact prediction Each evaluation network performs forward inference independently and outputs its own Q-value vector in parallel.
[0112] The long-term cumulative impact prediction values output by each actuator evaluation network are used to evaluate the long-term cumulative impact prediction values. The data is then sent back to the parameter server. The parameter server uses the Q value and a preset risk threshold as a reference. deviation ( For the current greedy action, calculate the policy gradient correction for each executor:
[0113] ;
[0114] in This is the policy learning rate (typically 0.001). For the parameters of the a-th executor policy network, the parameter server will... The policy network is broadcast to the local policy network of each actuator, and each policy network updates its parameters synchronously according to the gradient correction amount: replace This completes one distributed policy iteration. This process is repeated to achieve distributed incremental convergence of the optimal cooperative control policy, while ensuring global consistency in policy updates across all executors.
[0115] Through the above technical solution, after the risk-constrained reinforcement learning model adopts a distributed architecture, the independent evaluation networks corresponding to each actuator realize parallel inference, and the policy update achieves global consistency through the aggregation of global context feature tensors and the broadcasting of policy gradient corrections by the parameter server. Compared with the centralized single-point decision architecture, the decision delay of this distributed solution only increases linearly rather than quadratically when the scale of the actuators is doubled, which improves the real-time control capability in large-scale multi-point collaborative emission reduction scenarios.
[0116] To address the misalignment of control effects caused by the delayed action activation of global slow-response actuators and local fast-response actuators, the system introduces an "effect prediction and look-ahead alignment" mechanism in hierarchical decision-making.
[0117] When the global response agent set outputs a coarse-grained basic action Then, the collaborative control module will call a simplified effect prediction model based on the current environmental state (such as a box model based on mass conservation or an empirical transfer function) to quickly estimate the expected impact curve of the action on the pollution concentration in the key area over several local control cycles in the future (covering the response period of the local actuators).
[0118] Local compensation agent ensemble generates high-frequency disturbance cancellation actions At this time, the input state not only includes the current instantaneous fluctuation characteristics, but also incorporates the expected impact curve of the aforementioned global actions. Specifically, the policy network of the local agent uses the net fluctuation of "current measured fluctuation" minus "future expected improvement" as part of the input. This makes the goal of the local compensation action no longer to offset the current original fluctuation, but to offset the disturbances that remain after deducting the future effects of the global actions and still need to be dealt with urgently. This achieves alignment between macro-trend control and micro-instantaneous compensation in terms of temporal effects, avoiding control conflicts or overshoot.
[0119] In one embodiment, the system further includes a self-evolutionary correction module deployed in the central processing unit, dedicated to real-time monitoring of the prediction accuracy drift of the time series prediction module, and automatically triggering global retraining of the transition probability matrix when the deviation exceeds the allowable range.
[0120] The self-evolutionary correction module receives environmental feedback data (i.e., the measured pollution state observation sequence after the execution of control actions) from the instruction execution module in real time. Simultaneously, the pollution probability distribution prediction sequence for the corresponding time moment is retrieved from the time series prediction module. Transform the measured sequence into an empirical distribution. (That is, the frequency histogram of each state level, after normalization, satisfies) ) and predicted distribution (i.e., each time point) The average probability distribution within the decision window (also normalized) is compared.
[0121] The KL divergence between the empirical distribution and the predicted distribution is calculated using the following formula:
[0122] ;
[0123] To avoid The problem of time-logarithmic divergence affects the prediction of distributions. Perform Laplace smoothing: ,in Let j be the prediction frequency of state j within the decision window. Total frequency Smoothing coefficient (typical value) In the calculation of KL divergence, using... Alternative And agreed when When the value is 0, the contribution of this item is 0.
[0124] KL divergence value A larger value indicates a greater deviation between the measured and predicted distributions, meaning the transition probability matrix P's ability to fit the current pollution dynamics has decreased. KL calculations are performed in each fixed assessment period. (For example, execute once every 24 hours, or trigger immediately after each major emission event) Execute.
[0125] Compare the KL divergence value with a preset deviation threshold. (Values range from 0.05 to 0.2, set by system maintenance personnel based on prediction accuracy requirements) Compare. When KL > At this time, the self-evolutionary correction module sends a global retraining trigger command to the time series prediction module. The time series prediction module responds to this command by retrieving the most recent data from the historical database. Months (typical value) For historical environment samples (=3), re-perform maximum likelihood estimation, calculate the updated state transition probability matrix, and then smooth it (e.g., by adding Laplace smoothing to avoid the zero probability problem) to generate a new transition probability matrix. Replace the original matrix P and put it into use. When KL≤ At this time, the existing P remains unchanged, and normal operation continues. Through the self-evolution mechanism driven by the above quantitative criteria, the prediction accuracy of the system can continuously self-correct as the pollution scenario changes, avoiding linear degradation of model performance over time.
[0126] It should be clarified that the offline retraining of the transition probability matrix P and the online policy update of the risk-constrained reinforcement learning model adopt a hierarchical asynchronous operation mechanism:
[0127] Macro-evolution layer (offline retraining): Triggered by the self-evolution correction module, it performs global retraining on the state transition probability matrix P based on accumulated historical monitoring data, with a time scale of days or weeks, to capture the long-term trend changes and seasonal characteristics of pollution evolution.
[0128] Micro-adaptive layer (online update): Executed by the risk-constrained reinforcement learning model in the collaborative control module, with a time scale of seconds or minutes, the policy network is continuously updated online through the distributed iteration of policy gradient correction and parameter server, in response to short-term fluctuations such as wind speed, humidity, and emission intensity.
[0129] The two operating mechanisms described above are independent and do not conflict with each other. After the transition probability matrix P of the macro-layer is retrained, the new transition probability matrix will serve as the base model for the next stage of time series prediction, while the policy network update of the micro-layer will continue on this base model without interruption or reset. Through this layered decoupling design, the system simultaneously possesses the ability to track long-term evolutionary patterns and the ability to respond quickly to instantaneous disturbances.
[0130] To prevent a sudden increase in KL divergence due to a single pollution event, which could trigger unnecessary retraining and impair the model's generalization ability in normal scenarios, the self-evolutionary correction module employs the following enhancement logic:
[0131] 1. Persistence Criterion: The retraining trigger command is no longer based solely on the KL divergence value of a single evaluation, but requires that the moving average of the KL divergence continuously exceeds a threshold over N consecutive evaluation periods (e.g., N=3). This ensures that only persistent, systematic prediction biases trigger retraining, effectively filtering out transient anomalies.
[0132] 2. Dynamic Threshold Adjustment (Optional): The system can maintain a dynamic threshold table associated with the season or the dominant weather model. During seasons with naturally high forecast uncertainty (such as periods of frequent stable weather in winter), the system automatically adopts a higher threshold. During seasons with favorable diffusion conditions, a lower [specification rate] is adopted. The threshold can be determined by analyzing the reasonable fluctuation range of prediction error in different seasons in historical data.
[0133] 3. Training data filtering: When retraining is triggered, the time series prediction module will automatically remove the time period data marked as "extreme events" within the evaluation period when calling historical environmental samples (this label can be automatically generated by the risk quantification module based on the CVaR index), ensuring that the retraining focuses on learning the evolution law of normal pollution and improving the generalization performance of the new transition probability matrix.
[0134] Through the above technical solution, the statistical deviation between the measured pollution distribution and the predicted distribution of the transfer probability matrix is periodically tested by using KL divergence as a quantitative criterion, and the historical sample retraining is automatically triggered when the deviation exceeds the threshold. The prediction accuracy of the system can be automatically corrected according to the seasonal changes of pollution scenarios, the evolution of long-term emission patterns and extreme climate events. In the continuous 12-month tracking test, the prediction accuracy of the system equipped with the self-evolution correction module remained at a high level, while the prediction accuracy of the control system without the module degraded significantly over time.
[0135] This application further proposes that the instruction execution module includes priority preemption logic, which is used to forcibly interrupt the output of the regular control strategy and trigger emergency response actions for the peak risk area when the conditional risk value index exceeds a preset risk threshold.
[0136] The instruction execution module executes at a fixed period. (Typical value) =10 seconds) Real-time reception of the risk quantification module output. Value, and compared with the preset risk threshold. Compare them.
[0137] when > When the priority preemption logic triggers an interrupt signal, it forcibly suspends the instruction output queue of the current optimal cooperative control strategy π*, preventing further expansion of extreme risk exposure due to the execution delay of the regular strategy.
[0138] Extract the coordinates of the grid cell with the largest RF value within the current decision-making period from the dynamic risk field function RF(x,y,t+k). (i.e., the peak coordinates of the risk field), using these coordinates as the query key, retrieve data from the preset actuator location mapping table (an ordered list stored in the format <actuator ID, latitude and longitude coordinates, coverage radius>):
[0139] Where dist is the Euclidean distance (unit: meters). Let be the rated coverage radius (in meters) of the a-th actuator. If multiple actuators meet the coverage condition, then the actuator that covers the peak coordinates and has the highest overall control efficiency is retrieved, and its device identifier is returned. Effectiveness can be determined by a combination of factors such as distance, equipment type, power, and wind direction (or by weighting according to coverage radius).
[0140] in accordance with Exceeding The amplitude, to the actuator The corresponding hardware device sends one of two types of emergency commands: when Exceeding the severe threshold (Value range 0.2~0.5×) When ), send an emergency shutdown command (immediately close the valve of the nearest emission source, forcibly interrupting the emission); when Below At that time, an over-compensation pulse command is sent (to operate at the actuator's rated power limit for a short period, focusing on purifying the peak area). Emergency commands have higher priority than all regular control commands, until... Falling back to Following this, the system uses the current environmental feedback data to re-trigger the complete decision-making process from the data acquisition module to the collaborative control module, and calculates a new optimal collaborative control strategy. The system uses this new strategy as the regular control output. This design ensures that after an emergency exit, the system can make new decisions based on the latest environmental conditions, rather than simply restoring the strategy before the interruption, thus avoiding strategy failure caused by changes in environmental conditions during emergency actions.
[0141] Through the above technical solution, the introduction of priority preemption logic enables the instruction execution module to complete the entire process of routine strategy interruption, peak coordinate retrieval and emergency instruction sending within one monitoring cycle (typically 10 seconds) when the conditional risk value index exceeds the preset threshold. Compared with the traditional solution that relies on the next complete decision cycle to trigger emergency actions, the emergency response latency is reduced by about 80%, effectively preventing the expansion of pollution spread area and the ineffective idleness of emission reduction resources caused by response delays in the initial stage of extreme pollution events.
[0142] The following is a specific implementation of a smart environmental protection dynamic monitoring system based on multi-source data fusion:
[0143] Taking a chemical industrial park as an example, the park covers an area of approximately 2km × 3km and has 12 fixed gaseous pollutant concentration sensor stations (measuring SO2 and NO) within it. X The system comprises 10 hardware actuators: 8 meteorological sensors (monitoring pollutants such as PM2.5), 4 online continuous emission monitoring systems (CEMS) (corresponding to 4 main emission sources), 6 air purification devices, and 4 sprinkler systems. It is built based on multi-source data fusion to create a smart environmental dynamic monitoring system. Deployed at the park's edge computing node, the system collects raw data from the sensors in real time via a dual-link connection using ModBus / TCP and NB-IoT protocols. The data acquisition module uses a modified Z-score method to identify outlier observations for each sensor channel dimension-by-dimensionally. Outlier thresholds are set by calculating the historical mean and standard deviation of each channel. =3.0, marking observations exceeding the threshold as missing; calling the pre-trained GAN on the missing dimension, using the complete feature sequence of the non-missing dimension as conditional input to perform spatiotemporal completion on the missing position, generating a multi-source environmental data matrix with complete dimensions; then interpolating and aligning according to the time step Δt=1 minute, mapping to a 100m×100m spatial grid, and constructing a multi-dimensional environmental feature tensor T with a shape of 20×30×60×6.
[0144] The state evaluation module extracts SO2 and NO from tensor T. X The concentration characteristics of three pollutants, PM2.5, etc., are used to classify each pollutant into 6 discrete levels according to the national standard GB3095-2012, generating a discrete pollution state vector at the current moment. The size of the state space is =216 states. The time series prediction module calls the transition probability matrix P (216×216) trained by maximum likelihood estimation from the historical monitoring data of the past 3 years, as the pollution discrete state vector. Using the initial distribution, a pollution probability distribution sequence for the next 60 steps (1 hour) is generated through matrix exponentiation. ,..., Simultaneously, current wind speed, humidity, and real-time emission intensity from four emission sources are extracted to calculate the combined rate of change. Furthermore, the transition probability matrix is reweighted online with a modulation coefficient δ=0.3 to generate a dynamically corrected transition probability matrix P' to replace P for forecasting, enabling short-term forecast results to respond in real time to sudden weather changes.
[0145] The risk quantification module uses the coordinates of the emission sources where the four CEMS are located as the center, and calls a Gaussian kernel function to perform spatial diffusion mapping on the expected pollution level values at each time step in the pollution probability distribution sequence to generate a preliminary risk distribution. Load 30m resolution DEM data of the park from the terrain database, extract the slope θ(x,y) and shading coefficient ρ(x,y) of each grid cell, calculate the damping coefficient D(x,y) with shading attenuation coefficient β=1.0 and slope damping weight κ=0.5, and multiply R0 and D point by point to generate a dynamic risk field function. Perform tail integration on the RF with a confidence level of α=0.95 to obtain the conditional value-at-risk index. This serves as a constraint criterion for subsequent coordinated control.
[0146] The collaborative control module divides the 10 actuators into a global response agent set (4 purification devices with a response delay ≥ 60s) and a local compensation agent set (6 spray systems with a response delay < 60s) based on a 60-second response delay threshold. It then generates a coarse-grained emission reduction action subspace AG (4 levels for each of the 4 devices, totaling 256 combinations) and a high-frequency disturbance elimination action subspace AL (2 levels for each of the 6 spray systems, totaling 64 combinations). These subspaces are then processed using Cartesian product to generate a multi-agent collaborative action space A with a size of 256 × 64 = 16384. , The concatenated vectors (t, CVaRt) are used as the state input. Introduced as a Lagrange multiplier in the loss function, with the global reward function consisting of environmental quality improvement, energy consumption penalty, and risk penalty as the optimization objective, the algorithm achieves this through the collaborative iteration of the independent evaluation network of each executor and the parameter server within a distributed reinforcement learning architecture. Solve for the optimal cooperative control strategy π* under the constraint of (empirical calibration value).
[0147] The instruction execution module continuously receives commands in 10-second monitoring cycles. When it is below the threshold, π* is converted into control commands for each actuator and issued normally; when When the value is greater than 0.65, the priority preemption logic is triggered, forcibly interrupting the π* output. The system retrieves the current risk peak coordinates, locates the nearest actuator covering those coordinates in the actuator position mapping table, and sends an emergency shutdown or overcompensation pulse command based on the threshold amplitude, completing the emergency response within 10 seconds. After execution, the measured data from each sensor is transmitted back to the data acquisition module as environmental feedback data, triggering the rolling update of tensor T to form a closed-loop control loop. Simultaneously, the self-evolutionary correction module calculates the KL divergence between the measured and predicted pollution distributions every 24 hours. When KL > 0.1, it automatically triggers global retraining of the transition probability matrix to maintain prediction accuracy within an acceptable level.
[0148] The verification results of this embodiment show that during the 30-day continuous operation test period, the effective data coverage of the multidimensional environmental feature tensor reached 99.3%, the accuracy of pollution state prediction for 1 hour was 87.2%, the number of extreme pollution events (CVaR exceeding the threshold) decreased by 74%, the average emergency response latency was 8.3 seconds, the overall energy consumption was reduced by 22%, and the KL divergence triggered retraining 3 times without significant degradation in model prediction accuracy, thus verifying the engineering feasibility and technical effectiveness.
[0149] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. A smart environmental protection dynamic monitoring system based on multi-source data fusion, characterized in that, include: The data acquisition module is used to acquire multi-source environmental raw data of the target area and perform spatiotemporal synchronization and alignment to generate a multi-dimensional environmental feature tensor. The state assessment module is used to extract the concentration feature values of the multidimensional environmental feature tensor and map them to a preset pollution level range, and output the pollution discrete state vector at the current moment. The time-series prediction module is used to call the transition probability matrix obtained by training from historical environmental samples, and perform Markov iterative prediction on the pollution discrete state vector based on the transition probability matrix to generate a pollution probability distribution sequence within a preset decision period. The risk quantification module is used to extract geographic coordinate features from the multidimensional environmental feature tensor, perform spatiotemporal weighting by combining the expected values at each time point in the pollution probability distribution sequence, and generate a dynamic risk field function; the dynamic risk field function is then pre-set with a confidence level of... The tail integral operation generates the conditional value at risk index. The collaborative control module is used to receive the real-time state parameters of each actuator corresponding to the multi-dimensional environmental feature tensor, and combine the geographical location weights of each actuator to construct a multi-agent collaborative action space. The pollution discrete state vector, the dynamic risk field function, and the conditional risk value index are synchronously input into a preset risk constraint reinforcement learning model. Under the constraint that the conditional risk value index is lower than a preset risk threshold, the optimal cooperative control strategy is calculated based on the multi-agent cooperative action space. The instruction execution module is used to output the optimal cooperative control strategy, receive environmental feedback data after execution, and send the environmental feedback data back to the data acquisition module to update the multidimensional environmental feature tensor.
2. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, Before generating the multidimensional environment feature tensor, the following processing is also performed: Outlier observations in the original multi-source environmental data are identified and removed to obtain a cleaned environmental data sequence. A pre-trained generative adversarial network is invoked, and the missing dimensions are filled in with spatiotemporal features using the non-missing dimensions in the cleaned environmental data sequence. After generating a multi-source environmental data matrix with complete dimensions, the multi-dimensional environmental feature tensor is constructed.
3. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The time series prediction module is also used for: Extract non-pollution dimension features from the multidimensional environmental feature tensor, and calculate the joint change rate of wind speed, humidity and preset emission intensity in the non-pollution dimension features; The transition weights of each state in the transition probability matrix are reweighted online according to the joint rate of change to generate a dynamically corrected transition probability matrix. The original transition probability matrix is then replaced with the dynamically corrected transition probability matrix for iterative calculation.
4. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The process of generating the dynamic risk field function includes: Using the coordinates of each pollution source in the geographic coordinate features as the center point, the Gaussian kernel function is called to perform spatial dimensionality reduction mapping on the pollution probability distribution sequence to generate a preliminary risk distribution. The topological terrain features of the target area are retrieved from the preset terrain database. The damping coefficient is calculated based on the topological terrain features. The damping coefficient is then multiplied point by point with the preliminary risk distribution to generate the dynamic risk field function with physical diffusion semantics.
5. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The collaborative control module is also used for: The conditional value at risk index is introduced as a Lagrange multiplier into the loss function of the risk-constrained reinforcement learning model; The difference between the conditional value of risk index and the preset risk threshold is monitored. When the difference is lower than the preset fluctuation range, the exploration boundary of the multi-agent collaborative action space is automatically compressed to generate an action space with tightened constraints.
6. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The collaborative control module employs a hierarchical partitioning logic when constructing the multi-agent collaborative action space, specifically including: Based on the response parameters of each actuator, the actuators are divided into a global response agent set and a local compensation agent set; Extract the macroscopic gradient features of the dynamic risk field function, input the macroscopic gradient features into the global response agent set, and generate a coarse-grained emission reduction action subspace; Extract the instantaneous fluctuation features of the multidimensional environmental feature tensor, input the instantaneous fluctuation features into the local compensation agent set, and generate a high-frequency disturbance elimination action subspace; The coarse-grained emission reduction action subspace and the high-frequency disturbance elimination action subspace are combined using a Cartesian product operation to generate the multi-agent cooperative action space.
7. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, In the collaborative control module, the training and inference of the risk-constrained reinforcement learning model are based on the following global reward function, defined as: ; Where t is the discrete time step index, t=0,1,2,…,T, and T is the preset decision period. To extract the amount of environmental quality improvement from the environmental feedback data, To extract energy consumption from the real-time status parameters of the actuator, This refers to the conditional value-at-risk indicator. and These are the preset balance weight coefficients.
8. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The risk-constrained reinforcement learning model adopts a distributed architecture, and the collaborative control module specifically performs the following operations: The spatiotemporal gradient features of the dynamic risk field function are sent as global state variables to a preset parameter server, and the global context feature tensor aggregated by the parameter server based on the feedback from the global executor is received. The independent evaluation network corresponding to each executor is invoked, and the candidate actions in the global context feature tensor and the multi-agent cooperative action space are used as inputs to calculate the long-term cumulative impact prediction value of the candidate actions on the conditional risk value index. The long-term cumulative impact prediction value is sent back to the parameter server. The parameter server calculates the policy gradient correction amount based on the deviation between the long-term cumulative impact prediction value and the preset risk threshold, and synchronizes the policy gradient correction amount to the policy network of each actuator to achieve distributed iteration of the optimal collaborative control policy.
9. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The system also includes a self-evolution correction module, used for: Receive the environmental feedback data returned by the instruction execution module, and call the pollution probability distribution sequence from the time series prediction module; Calculate the KL divergence value between the environmental feedback data and the pollution probability distribution sequence, and compare the KL divergence value with a preset deviation threshold; When the KL divergence value exceeds the preset deviation threshold, a global retraining trigger instruction is sent to the time series prediction module. The time series prediction module then re-invokes historical environment samples from the historical database to train the transition probability matrix according to the trigger instruction.
10. The intelligent environmental protection dynamic monitoring system based on multi-source data fusion according to claim 1, characterized in that, The instruction execution module also includes priority preemption logic, specifically including: The conditional risk value index generated by the risk quantification module is received in real time, and the conditional risk value index is compared with the preset risk threshold. When the conditional risk value index exceeds the preset risk threshold, the output of the optimal collaborative control strategy is forcibly interrupted, and the actuator identifier is retrieved from the preset actuator position mapping table according to the peak coordinate of the dynamic risk field function. An emergency shutdown command or an overcompensation pulse command is sent to the hardware actuator corresponding to the actuator identifier.